doctor

Preconfig Doctor Reads a Failed Agent Setup, Names the Cause and Writes the Fix Into the Spec

When the Alpha ran orders-api’s setup with Redis left out of the spec, the machine printed 1,096 lines. The setup itself worked: every package installed and PostgreSQL started. Then three tests failed, and the reason sat at line 778, between stack traces: a refused connection to port 6379. Someone has to find that line, work out which part of the setup it points to, and change the right file. On an agent platform, that someone can be a developer coming back to a half-finished task.

Preconfig Doctor does that reading. preconfig doctor takes the log, finds the step that broke and the lines that show why, and names the cause: the code connects to Redis on localhost:6379, and nothing is listening there, because the spec doesn’t run Redis. The fix is one line, Redis 7 under services. With --fix, Doctor writes it into preconfig.yaml, keeping the file’s comments, and every platform’s file is rebuilt. The same setup, run again on a clean Ubuntu 24.04 machine, was READY in 65.1 seconds. That’s the project’s premise at work: an agent’s machine should be ready before the agent starts, described once in a spec a person can review and proven on a clean machine by the project’s own tests.

How Doctor reads a log: the log’s steps, the first real error in the step that failed, 31 rules tried in order, and the answer: the cause and, when it’s in the spec, the change. Below, the Alpha’s recorded run without Redis, what Doctor prints for it, and the line –fix adds to preconfig.yaml.

Rules, and “Unknown” When None Fits

There’s no model inside. Doctor has 31 rules, each tying an error format to its cause: apt’s “Unable to locate package”, libpq’s refused connection, pip’s “requires a different Python”. It reads seven kinds of log, from the setup script’s own output to a Docker build. When no rule fits, it says “unknown”, shows the lines that look like errors, and changes nothing.

It also knows where a fix belongs. A missing service, package, runtime or variable is a change to the spec, and Doctor writes it. A blocked download, a stale lockfile, an unset secret or a test that fails on its own assertion gets an explanation instead, and the place to fix it.

How It Was Measured

The rules were tuned on 28 recorded cases, each run for real: 25 failures, nearly all planted on purpose in sample repositories, and 3 setups that passed. Then the rules were frozen, and 11 more cases were recorded, designed along with the rules but run only after the freeze. On those 11, Doctor was right 8 times and said “unknown” twice, for lxml and Pillow built from source, which word their missing headers their own way. Once it blamed the network correctly but didn’t name the package index it couldn’t reach. In 74 answers across both sets, it never wrote a wrong change.

A Fix That Uncovers the Next Failure

Every change Doctor wrote was run again on a clean machine, and every failure it named was gone. One case shows why that matters. psycopg2, built from source on a bare machine, first needed a C compiler. With that added, the build got further and stopped on Python’s headers. With those, it stopped on libpq’s. Doctor read each new log the same way, and the third fix reached READY. All 20 cases with a change got past their failure: 19 setups reached READY, and Cursor’s Dockerfile built.

What’s Next

Doctor’s logs so far came from one machine. The Beta gives it the platforms’ own failures from Copilot, Cursor and Codespaces runs, fixes the three misses first, and records fresh cases after every change to the rules, so the score stays honest. When verify fails on a pull request, the planned GitHub Action will carry Doctor’s cause and fix in its comment.

You can try it now in the Doctor demo: nine recorded failures, or a log of your own, read by the engine running in your browser. The Doctor page has the full results.