Game Anatomy

Controlled autopsies of AI game-building systems.

The Mechanic Was Real. Human Review Found It Too Subtle.

The code contained the mechanic. The automated gates could reach it. The build stayed healthy while they did.

Then one operator played it and barely felt the difference.

That is not a contradiction. It is the gap between proving that a mechanic exists and proving that it changes the play loop at a useful magnitude.

Command Cells passed the named gates

A Codex game-design-ideator selected Command Cells; Gemini Flash implementation produced the graft.

The concept made command ships part of the formation. Destroying one was meant to trigger a visible response in its column. The implementation exposed the required state and preserved the test surface.

Command Cells passed tier1, space_invaders_driver.mjs, and UX smoke, but one operator review found the mechanic source-real and play-weak.

Those are three named machine gates and one bounded human-terminal judgment. The machine verdicts say the artifact loaded, exposed the expected behavior, and remained operable. They do not say that the mechanic created enough pressure to matter during play.

The mechanic was source-real

The important first check was not a screenshot. It was the source and state transition behind the effect: command ships existed as a distinct role, hits could identify them, and their destruction changed the corresponding column.

Play Command Cells and look for the command-ship response inside the formation.

Command Cells Space Invaders artifact with the selected mechanic implemented in the formation.
SI08-C1 and SI08-C2: the concept came from the Codex game-design-ideator, while the graft was a Gemini Flash implementation.

This clears a useful evidentiary floor. The result was not a prompt label pasted onto an unchanged game. Something in the rules had changed.

One operator found it too subtle

The human-terminal question was different: did that rule change become legible soon enough, and strongly enough, to alter moment-to-moment decisions?

In one operator review, it did not. Command ships blended into a familiar firing loop. Their effect existed, but the player could continue with almost the same positioning, target priorities, and survival plan.

That judgment does not describe a player population. It records one operator's comparison of these artifacts under the lab's terminal review. Pressure, fairness, atmosphere, and replay desire remained outside the automated driver.

The earlier direct overlays had produced useful parameter tuning in these runs—more audiovisual pressure and adjusted presentation—but they had not supplied a play-visible rule change of the desired magnitude. Command Cells moved into rule space, yet still landed below the practical bar.

Set a 30-second magnitude bar

The next search used a practical test: within about 30 seconds, the mechanic should change at least one of these surfaces strongly enough for the operator to notice and respond:

This was a search rubric, not a new automated PASS gate. It forced the prompt to name the size of the desired change instead of asking vaguely for more fun, tension, or rhythm.

Falling Wreckage changed the survival geometry

Falling Wreckage made destroyed invaders leave debris. The debris could block shots, damage shields, and threaten the player. That created moving hazards between the formation and the cannon, so the effect entered both projectile paths and survival geometry.

Two implementations were kept separate. Despite a misleading lab filename, one was executed by a Codex-local operator. The other was executed by a Codex subagent. Neither execution came from Flash or Gemini.

Play the Codex-local operator version and compare its debris lanes with Command Cells.

Falling Wreckage implementation executed by a Codex-local operator, showing debris in the play field.
Codex-local operator execution: the filename is not executor evidence.

Play the Codex subagent version to inspect the second independent execution.

Falling Wreckage implementation executed by a Codex subagent, showing play-visible falling debris.
Codex subagent execution: a second implementation of the same higher-magnitude mechanic.

Compare the bounded operator judgment

In one operator review, the two Codex-local Falling Wreckage implementations were stronger than Command Cells and closer to the prior magnitude bar; neither was a Flash/Gemini execution.

The local-operator version was slightly preferred in that review. The useful comparison is not a model ranking. It is that debris changed where the operator could stand, when shots could pass, and how shields absorbed risk, while Command Cells left more of the familiar loop intact.

The comparison stays bounded to these artifacts and this reviewer. It does not establish universal preference, fairness, or long-session replay value.

Raise the bar without inventing a law

The separated-ideation candidate remained play-weak even though its mechanic was real and its machine gates were green. Falling Wreckage then produced a stronger bounded result after the search demanded a visible change within 30 seconds.

That sequence motivated a higher mechanic-magnitude bar. It did not isolate ideation structure as the cause, and it does not establish separation as a general requirement. The prompts, concepts, executors, and implementations changed together.

The operational lesson is narrower: verify that the rule exists, then ask whether it changes play soon enough to matter. If the second answer is still uncertain, the next experiment needs a magnitude criterion—not a stronger adjective.

The Space Invaders experiment keeps the full mechanics search beside all fifteen playable artifacts.