Preconfig Doctor

When an agent’s machine fails to set up, the log is long and the cause is buried. The Alpha’s own run with Redis left out of the spec wrote 1,096 lines, and the line that says why is line 778, a refused connection to port 6379 in the middle of a test runner’s stack traces. Preconfig Doctor reads that log for you. It finds the step that broke and the lines that show why, names the cause, and when the fix belongs in preconfig.yaml, it writes the change. Every platform’s file is then rebuilt from the spec, and verify proves the fix on a clean machine.

The project’s premise is that an agent’s machine should be ready before the agent starts, described once in a spec a person can review and proven on a clean machine by the project’s own tests. Doctor covers the case where that proof fails. It is the second engine of Preconfiguration, after Preconfig Core, and a command in the same program: preconfig doctor. It has no model and needs no service. Every answer comes from rules, and the same log always gets the same answer.

Where it stands. On 11 cases recorded after its rules were frozen, Doctor was right 8 times, said “unknown” twice and was scored wrong once, for naming the network without the host. In all 74 answers, tuning and holdout together, it never wrote a wrong change. Every case it wrote a fix for got past its failure when it ran again on a clean machine. Next it needs real logs from the agent platforms themselves. Try the Doctor demo

How Doctor reads a log: it reads the log’s steps, finds the step that failed and the first real error in it, and tries 31 rules in order. A rule that fits gives the cause and, when the fix belongs in the spec, the change. Below, the Alpha’s recorded run without Redis, what Doctor prints for it, and the line –fix adds to preconfig.yaml.

How It Works

  • Read. Doctor tells seven kinds of log apart by their markers: the setup script’s own output, verify’s terminal output and its events, a Docker build, a GitHub Actions log, cloud-init’s log, and plain text. It finds the steps and the one that failed.
  • Find. Inside the failed step, it looks past what a failure drags behind it: retries, stack frames, the test runner’s summary and its warnings.
  • Match. 31 rules are tried in order, the most specific first. Each ties a known error format to a cause and to where the fix belongs: the spec, the network, the repository, a secret, the code or the machine. When none fits, Doctor says “unknown” and shows the lines that look like errors. It never changes the spec on a guess.
  • Fix. For most causes in the spec, Doctor writes one of eight kinds of change: add or replace a package, add a service, set a runtime, add a tool, add or set a variable, add a secret. A spec that doesn’t load, or a variable whose value Doctor can’t know, gets an explanation instead. With --fix it edits preconfig.yaml line by line, keeping its comments, checks that the result still loads, and rebuilds every platform’s file.
preconfig doctor --dir REPO LOG          explain a failed setup from its log
preconfig doctor --dir REPO --fix LOG    and write the fix into the spec

Passwords in URLs, tokens, keys and authorization headers are masked in the lines Doctor quotes. It reads logs on the machine it runs on, so a log never leaves it.

How It Was Measured

Each case copies a sample repository, usually changes one thing, and runs the setup on a clean Ubuntu 24.04 container, or builds Cursor’s Dockerfile, keeping everything the machine printed: Redis left out of the spec, a misspelled package, a test that needs a locale, a source the network blocks. Each case’s expected answer was written when the case was designed, before its log existed, and corrected only from what its log showed. The rules were tuned on 28 cases, then frozen, and 11 more cases were recorded only after that. The score on those 11 is the honest one, with one caveat: the holdout cases were designed along with the rules, so they aren’t blind to them.

Set Logs Right Unknown Wrong
Tuning, full logs 25 25 0 0
Tuning, verify’s short terminal output 25 23 2 0
Tuning, Docker builds 3 3 0 0
Holdout, full logs 10 7 2 1
Holdout, verify’s short terminal output 10 7 2 1
Holdout, the spec’s own errors 1 1 0 0

The holdout found three things the rules hadn’t met. lxml and Pillow, built from source, print their own messages about missing headers, and Doctor said “unknown” on both. For a package index the network couldn’t resolve, Doctor named the network as the cause but not the host, which pip prints on an earlier line, and the scorer counts that as wrong. None of the three changed the spec. Each is a small rule for the next version.

Fixes Proven on a Clean Machine

Every change Doctor wrote was run again: the fix applied with preconfig doctor --fix, then the setup on a clean container, or Cursor’s Dockerfile built again. 20 cases had a change, 14 from the tuning cases and 6 from the holdout. In 22 runs, each failure Doctor named was gone: 19 setups reached READY, in 57 to 106 seconds each, Cursor’s Dockerfile built, and two runs of the psycopg2 case got further and stopped on the next missing package.

One case shows what a planted failure can hide. psycopg2, built from source on a bare machine, needed a C compiler, then Python’s headers, then libpq’s, and each run stopped on the next one. Doctor read each new log the same way, and the third fix reached READY.

Tests and Speed

Measure Result
Test functions 15, all passing, with 42 rule cases and 13 edit cases
Planted bugs 14 of 14 caught, after the first run showed two missing checks
Fuzzing, final round 643,585 logs and 2,116,863 spec edits, no failures
Coverage 84.1% of statements
Package names Doctor can write 106 of 106 found in Ubuntu 24.04’s archive
Time on the Alpha’s 1,096-line log About 30 ms in Go; 182 ms for the browser build, timed in Node.js 22
Browser build 1.2 MB, 438 KB compressed, with the same answers as the command line on all 112 recorded files

What It Doesn’t Do Yet

  • Real logs from the platforms. Every log so far was recorded on one machine, with causes nearly always planted on purpose. Copilot’s, Cursor’s and Codespaces’ own setup logs come with the Beta.
  • Some of its readers. GitHub Actions, cloud-init and plain-text logs were tested on written examples only, and verify’s events on one recorded run.
  • Other systems. Ubuntu 24.04 only, and only English error messages.

The Doctor demo on the Alpha’s recorded run without Redis: the log, the cause Doctor names from two of its 1,096 lines, the fix for preconfig.yaml, and the proof, READY in 65.1 seconds on a clean machine.

The demo runs Doctor in your browser on nine recorded failures, or on a log you paste: the Doctor demo. The full method and every figure are in the Doctor Alpha report, available on request through the contact page.