
1972
The museum piece
Claude Haiku
Two paddles, two human players, minimal classic Pong.
Controlled autopsies of AI game-building systems.
Evidence-led experiments on AI game generation, testing, and the verification harnesses that make results inspectable.
Subscribe by emailRead the boundary map
Every claim here is attached to a build you can run in your own browser. Start with the debut.
Same one-line spec, three different AI systems, three different interpretations.

1972
Claude Haiku
Two paddles, two human players, minimal classic Pong.

The Textbook
Gemini Pro 3.1
Mouse control, computer opponent, clean web-tutorial style.

The Casino
Gemini Flash 3.5
Power-ups, particles, synthesized sound, and extra spectacle.
Seven AI systems each got the same two-sentence prompt with nothing specified. All 21 artifacts came back as canonical templates, and the instrument turned out to be contaminated before the data was read.
One Space Invaders route received primitive FAIL, engineered mechanical PASS, and human-terminal practical FAIL—and all three verdicts were honest.
Command Cells passed every named machine gate, but one operator review found the mechanic too subtle and set a 30-second magnitude bar for the next search.
The accepted 03.2B Space Invaders base combined a cleaner GPT seed, a Flash presentation donor, and a Pro executor under machine and human verification.
In a Breakout experiment, mechanically testable builds passed while uninstrumented grafts won the human review.
All three engineered Breakouts passed the mechanical driver, but none provided the player a complete restart loop.
A deterministic Pong test harness falsely failed correct AI-generated games twice, showing why falsifiers need their own autopsy.
An AI Pong specification can enforce correctness, but this experiment shows why player feel still needs separate human judgment.
Three AI systems turned one vague Pong prompt into different games, then an explicit test contract made the results verifiable.
Each rung of the ladder is one game of known difficulty, the same models, and an explicit contract that decides what counts as done.
A pre-registered probe on what seven model-systems put into a game prompt that specifies nothing: 21 artifacts, blind human review, and zero distinctive mechanical cores.
A controlled Pac-Man experiment on autonomous enemies: practical completeness collapsed from 3/4 to 1/4 while the technical floor barely moved.
A controlled Space Invaders experiment separating sparse-prompt behavior, mechanical contracts, model roles, and human gameplay judgment.
A controlled Breakout experiment separating mechanical verifiability from the human-facing quality of an AI-generated game.
A controlled Pong experiment on how explicit specifications and deterministic test APIs make AI-generated games verifiable.
New autopsies by email. No schedule promises, no filler.