DRAGONS CRACK

Devlog

Four days of an agent auditing the game

On the morning of 24 September the owner gave an autonomous coding agent, OpenAI’s Codex, one instruction: audit the whole game, find bugs, and keep going until its usage ran out. It ran for four days and stopped mid-pass on the 27th when its credits did. This is what it did, where it fell short, and what closing it out took.

What it fixed

About sixty scoped repairs, each with the same shape: reproduce the fault against the old code, install a narrow fix, write a regression, and run the modules the change touched. Save slots and corrupt saves. Replay drift from rounded stores. Eight save-and-load replay cases. Victory retry. The Rising’s credit. The raid camera’s sightline through a sagging roof. Input leaking through overlays. A rebound key that lost its binding on restart. Smelting output that vanished, and construction offered during a defense. It also cut 42 short gameplay reels for a phone, and built the music and the interface pass that the previous post describes.

Where it fell short

Every one of its sixty checkpoints ran a selected subset of the tests, between 91 and 507, and no run of the whole registry ever completed on the tree it left. Its three attempts at the full replay matrix each hit a two-hour deadline; one of them was paused 1,946 times by the thermal guard, for 4,956 seconds in all, because this laptop reaches its critical temperature within minutes under load. The registry’s own hygiene test failed on its tree: two new test modules had no main guard. The audio dropout investigation ran on a software renderer at a frame a second and fixed the wrong thing. Two correct camera repairs to one code path together scanned 537,538 triangles a frame for the rest of every fire. And several findings were enshrined by tests written beside the repair: a drake that re-broke off on every later hit, a dawn that walked sheltering locals out into a visit, a bell field that waited for the next visit, a pilot that could not make room for water with a stone-filled pack.

Closing it out

The close-out ran the whole registry, 115 modules, from a copy of the tree under a virtual display with the same thermal guard: 1,965 tests, two failures, no errors, nine skips, in 4,224 seconds. The two failures were one expectation the audio repair had changed; the nine skips read older save formats from git history the copy did not carry. The replay matrix, twenty traces by three frame schedules plus save and load, passed inside that run in 2,477 seconds, the gate the agent’s three attempts never reached. The eighteen modules changed after the copy were rerun in the tree: 625 tests, nothing failed or skipped. Every one of the ten characters and the raid was then played by the autopilot on the laptop’s real display and exited cleanly. The mixed audio bus was tapped on that display and, from a cooled package, showed no silence of ten milliseconds or more after the first sound.

Three candidates were left on purpose for the owner: a rebuilt Eastern Lung whose legs no longer share a flank, the material relighting behind it, and a dropped shadow certificate. The lessons went into the project’s learnings file as its nineteenth section. The shortest of them: a ledger that says N selected tests passed is not a suite, a software renderer is not the player’s machine, and a run that was paused is not a timing result.

Sources

Where this comes from

Every fact and number above is taken from these records in the game’s own repository, which is not public.

  • the agent’s own packet: each repair, its evidence and the replay matrix that hit its deadline
  • the close-out ledger: the gates, the shortfalls lane by lane, the repairs and the final runs
  • the full-registry run, the in-tree rerun, the real-display play checks and the audio taps
  • section 19: closing out an agent’s four-day run

Every post, and the devlog’s RSS feed.