JevRLREWARD LABGitHub

TYPESAFE JEV × REINFORCEMENT LEARNING

Reward by JEV.
Learn by playing.

Fast, typesafe, and accurate.
JEV supplies the reward. The agent learns to play four classic games.

01 → 04

Same games. Same DQN.
Different reward signals.

Explore the evidence ↓
JEV supplies the reward: environment-agent-JEV learning loop alongside measured CartPole Total score curves for JEV, Human design, and Native.
JEV supplies the training reward. CartPole total score is the game’s original cumulative score. Curves show mean ± SD across 3 training seeds, with 20 development episodes per checkpoint.

EXPLORE THE EXPERIMENT

Choose a game. Train your own.

Inspect every checkpoint, compare reward sources,
or start a fresh local training run.

CLASSIC CONTROL

CartPole

JEV REWARD

LOADING RECORDED EXPERIMENTSTEP 0
FIXED REPLAY SEED 10000
Total score —
Training time machineStep 0 · episode 0
Evaluation —

MEASURE WHAT THE AGENT ACTUALLY ACHIEVES

The reward is a signal.
The game is the test.

Training scores have different scales.
Compare policies using total score and success rate.
Shaded bands show variation across training seeds.

Total score

Original cumulative game score · development evaluation · mean ± training-seed SD

Success rate

Development evaluation · mean ± training-seed SD

Training reward

Selected run · per-episode return and a 20-episode moving mean

● JEV● Native● Human design

Protocol: 3 training seeds · equal step budgets within each game · 20 development episodes per checkpoint · 100 disjoint final test episodes. Observations and rubrics are human-designed. These experiments measure a reward integration, not a new RL algorithm or a claim of general game intelligence.