← Rocket CabBuilding with Astra · Eric HubbardPublished September 27, 2026
Rocket Cab / AI training

Training AI to play
Rocket Cab.

Four policies, 100 million training steps each. The clearest finding came from watching how one agent earned its points.

The question

Would scoring rewards encourage sustained pinball play?

The observation

One final policy earned most of its score from three opening sequences.

The next test

Measure ball retention and launch progress in unrestricted games.

An award. An abort. Another opening.

The final control-discipline policy completed ARM, IGN and REL from the opening plunge, then triggered a hard abort. The next ball repeated the pattern. Three launch awards supplied 175,000 of its 187,500 points—about 93% of the score.

  1. Ball 1 · 4.97s+25,000Abort at 5.07s
  2. Ball 2 · 12.02s+50,000Abort at 12.15s
  3. Ball 3 · 19.10s+100,000Abort at 22.62s
Recorded agent gameplayOriginal physics
Rocket Cab pinball cabinet; press Play to watch recorded agent gameplay.
SCORE0
GAME TIME0.0s
BALL1/3
Compare all eight recorded games

Enable JavaScript to play this recording.

Press Play to watch the first award and abort at ¼ speed. Recorded final checkpoint, control-discipline run. Playback uses the original physics and verifies the recorded result at the end. This is one observed game, not an estimate of typical performance. No leaderboard submissions.

More training did not mean steady improvement.

We compared the original score reward with a longer reward horizon, removal of drain-bonus points from the reward, and small costs for rapid paddle re-presses and nudging. Holding a paddle remained free. The physics and displayed game score were unchanged.

Four evaluation score histories through 100 million training steps, showing fluctuations and plateaus rather than steady gains
Rolling means of up to 20 deterministic evaluations at successive checkpoints. One training seed per variant; missing earlier history is left blank. These are exploratory runs, not independent repeated trials.

The longer-horizon run finished with the highest recent mean: 244,349, compared with 192,852 for the reference. Three variants reached full liftoff in 3 of their final 20 evaluations; the reference reached none. Those results identify follow-up questions, rather than establish a reliably superior policy.

A learned policy can still favor a narrow pattern.

The agent received ball position and velocity, paddle state, targets and timers. Training sampled actions for exploration; these recorded evaluations use deterministic choices from the learned policy. The inputs were not simply random button presses.

But game-state observations do not guarantee controlled shots. The fixed, full-charge opening offered repeatable points, while the hard-abort penalty was small relative to the launch awards. The observed sequence is consistent with collecting opening rewards and spending balls to repeat them. It does not prove deliberate planning.

The fourth launch stage is worth 250,000 points. Reaching it requires more than the three convenient openings. Some best checkpoints achieved it; this final checkpoint did not. A selected highlight and the policy at the end of training tell different stories.

Measure progress beyond the opening.

The follow-up compares five fresh policies without a game clock, including a stronger hard-abort penalty. Launches per ball, abort frequency and returns after paddle contact will help distinguish point collection from sustained play.

That series was still running at this September 27, 2026 snapshot. Its results remain open. The useful lesson so far is methodological: score curves show whether performance changes; recordings help explain what behavior produced those scores.

Experiment settings, full results & limitations
Completed capped runs · last 20 checkpoint evaluations
PolicyMean scoreBest observed*Full liftoffs
Reference192,852480,3400 / 20
Longer horizon244,349553,2403 / 20
No drain bonus237,838553,4103 / 20
Control discipline237,690513,3603 / 20

*Best observed across each run, selected by evaluation score. The final 20 samples use successive policy snapshots, not 20 independent trials of one frozen policy. Each variant used one seed; the original three include checkpoint resumes. Whole PPO rollouts slightly exceed the 100-million-step targets.

PPO (proximal policy optimization) updated a shared policy from parallel games. Decisions occurred 60 times per simulated second. The reference discount factor was 0.99; the longer-horizon variant used 0.999, giving later rewards more weight.

The control variant charged 0.05 reward units for an accepted nudge, plus the ordinary 1-unit drain cost when an abort lost the ball. The first three launch awards supplied 25, 50 and 100 reward units. The follow-up adds 100 units of hard-abort cost. These incentives do not deduct from the displayed game score.

Across the full histories, time limits ended 0 reference, 65 longer-horizon, 95 no-drain-bonus and 10 control-discipline evaluations. Zero timeouts in the final 20 does not exclude an earlier effect on learning. Human games have three balls and no clock, so their scores are excluded from this comparison.

The new series uses 12 parallel games per policy, 60 total, and independent evaluations of frozen checkpoints. Unfinished evaluations contribute no completed score. Update settings also changed; differences between series cannot be attributed solely to removing the clock.

Download the evidence