An award. An abort. Another opening.
The final control-discipline policy completed ARM, IGN and REL from the opening plunge, then triggered a hard abort. The next ball repeated the pattern. Three launch awards supplied 175,000 of its 187,500 points—about 93% of the score.
- Ball 1 · 4.97s+25,000Abort at 5.07s
- Ball 2 · 12.02s+50,000Abort at 12.15s
- Ball 3 · 19.10s+100,000Abort at 22.62s

Compare all eight recorded games
Enable JavaScript to play this recording.
Press Play to watch the first award and abort at ¼ speed. Recorded final checkpoint, control-discipline run. Playback uses the original physics and verifies the recorded result at the end. This is one observed game, not an estimate of typical performance. No leaderboard submissions.
More training did not mean steady improvement.
We compared the original score reward with a longer reward horizon, removal of drain-bonus points from the reward, and small costs for rapid paddle re-presses and nudging. Holding a paddle remained free. The physics and displayed game score were unchanged.

The longer-horizon run finished with the highest recent mean: 244,349, compared with 192,852 for the reference. Three variants reached full liftoff in 3 of their final 20 evaluations; the reference reached none. Those results identify follow-up questions, rather than establish a reliably superior policy.
A learned policy can still favor a narrow pattern.
The agent received ball position and velocity, paddle state, targets and timers. Training sampled actions for exploration; these recorded evaluations use deterministic choices from the learned policy. The inputs were not simply random button presses.
But game-state observations do not guarantee controlled shots. The fixed, full-charge opening offered repeatable points, while the hard-abort penalty was small relative to the launch awards. The observed sequence is consistent with collecting opening rewards and spending balls to repeat them. It does not prove deliberate planning.
The fourth launch stage is worth 250,000 points. Reaching it requires more than the three convenient openings. Some best checkpoints achieved it; this final checkpoint did not. A selected highlight and the policy at the end of training tell different stories.
Measure progress beyond the opening.
The follow-up compares five fresh policies without a game clock, including a stronger hard-abort penalty. Launches per ball, abort frequency and returns after paddle contact will help distinguish point collection from sustained play.
That series was still running at this September 27, 2026 snapshot. Its results remain open. The useful lesson so far is methodological: score curves show whether performance changes; recordings help explain what behavior produced those scores.
Experiment settings, full results & limitations
| Policy | Mean score | Best observed* | Full liftoffs |
|---|---|---|---|
| Reference | 192,852 | 480,340 | 0 / 20 |
| Longer horizon | 244,349 | 553,240 | 3 / 20 |
| No drain bonus | 237,838 | 553,410 | 3 / 20 |
| Control discipline | 237,690 | 513,360 | 3 / 20 |
*Best observed across each run, selected by evaluation score. The final 20 samples use successive policy snapshots, not 20 independent trials of one frozen policy. Each variant used one seed; the original three include checkpoint resumes. Whole PPO rollouts slightly exceed the 100-million-step targets.
PPO (proximal policy optimization) updated a shared policy from parallel games. Decisions occurred 60 times per simulated second. The reference discount factor was 0.99; the longer-horizon variant used 0.999, giving later rewards more weight.
The control variant charged 0.05 reward units for an accepted nudge, plus the ordinary 1-unit drain cost when an abort lost the ball. The first three launch awards supplied 25, 50 and 100 reward units. The follow-up adds 100 units of hard-abort cost. These incentives do not deduct from the displayed game score.
Across the full histories, time limits ended 0 reference, 65 longer-horizon, 95 no-drain-bonus and 10 control-discipline evaluations. Zero timeouts in the final 20 does not exclude an earlier effect on learning. Human games have three balls and no clock, so their scores are excluded from this comparison.
The new series uses 12 parallel games per policy, 60 total, and independent evaluations of frozen checkpoints. Unfinished evaluations contribute no completed score. Update settings also changed; differences between series cannot be attributed solely to removing the clock.
Download the evidence
- Completed experiment snapshot — settings, histories and checkpoint hashes.
- Recorded event timings — launches, drains and aborts across eight replays.
- Featured recording — the 187,500-point final game.
- Longer-horizon highlight — a selected 553,240-point game.
- Worker benchmark — local throughput pilot and its conditions.