← Field notes

Plan first, then prove it

How a writing run in muxcode plans a change, builds it and proves it, and why a model's yes is never enough to pass the work.

A writing run in muxcode, one that changes your code, plans first, builds second and proves it last. No model’s yes passes the work alone: the host, muxcode’s own code around the models, checks the change before any judge does, and no model can pass failed verification.

The seats, in order

A run has four seats by default, each a role with its own model and effort. Effort is how hard a model reasons: low, medium, high, xhigh or max.

  • Root: GPT-6 Luna, the model you chat with, starts the run.
  • Architect: Claude Opus 5.5, fixed at high, plans, reviews, replans and judges.
  • Worker: Claude Opus 5.5, medium up to xhigh, writes the change.
  • Retry: Claude Fable 5.1, xhigh up to max, takes over after a failure.

Jev, the classifier, is TypeSafe’s System One model: it answers the host’s questions with probabilities, which the host holds to fixed bars. It runs on a mux Pro or Team plan, or with your own TypeSafe key. Before most worker and retry calls, Jev picks a level in the seat’s range; the host applies picks at 0.85 or higher. In the first live run with this worker, filed 2026-09-24, Jev kept it at medium all 4 times. We don’t know yet how often it raises one.

No other seat reviews the plan: the Architect re-reads its own in a fresh session. The judge shares the worker’s model, so only a retry brings in a second model, Fable. We claim on our homepage that Jev picks the cheapest worker; with one, there’s nothing to pick. In Settings → Agents, add a second worker for work the Architect marks routine, or give the worker a model the judge doesn’t use.

The contract

The Architect’s plan is one to four contracts, written limits a worker stays inside, each a route worked in turn. Each names the goal, the kind of change, the paths the worker may edit, frozen ones it may not, and the commands that verify it. A context gap, something only you can supply, stops the run.

Before the worker starts

The host screens the plan before any worker runs.

  • A mechanical screen catches a check you named in backticks that nothing runs, and verification that can’t fail, like npm test || true.
  • If that finds nothing, Jev reads whether a contract skips a named check, needs a file it can’t edit, breaks your rule, or would pass a wrong change. A yes at 0.7 sends it back. Both screens share one send-back per run.
  • A lint, the host’s static check, sends back a plan naming a missing program or a folder that doesn’t exist; a second failure stops the run.

Then the host runs the contract’s checks and up to 10 of the project’s own, none that write files, on the untouched tree: the baseline.

While the worker works

The host watches the worker and can steer or end it. After each edit, it runs any per-file checks the plan names and shows the worker what broke. Three identical failures of one command earn a steer, a host message to change course; five end the worker. Every 8 tool calls, Jev reads whether the worker has drifted or is stuck: one sure reading at 0.85 steers it, a second ends it.

A worker that reports itself blocked goes to triage, where Jev reads why. A 0.8 reading of a missing tool or a question only you can answer pauses the run, retry unspent. Without Jev, only a missing program the worker names does.

The gate, then the judge

The gate, the host’s checks on the finished attempt, costs no tokens; the Architect judges only after it passes.

  • A file changed outside the contract pauses the run; a frozen path touched stops it.
  • Every verification command must exit 0. A secret scan fails a known key format the change adds, unless the worker marks its line as a fixture, which the judge sees.
  • A project check counts only if it fails twice. It then fails the change if it passed before, or shows a new failing test, more failures, or a new failure Jev reads at 0.85.
  • A rewritten script that checks less, or an added test that only checks the code ran, fails on a 0.85 Jev reading. Without Jev, code catches only a script that can’t fail or is gone, or a test asserting nothing or only that something exists. Any Jev answer replaces that rule, so one under 0.85 passes a test the code would fail.

The Architect judges through four lenses, each pass or fail: the contract, the tests, the scope, and your own words, which every seat reads. An accept that fails a lens is a reject. A reject goes to the Fable retry, replanned first if the plan was at fault, once per route; a second failure stops the run.

What the shape really is

A writing run is a chain, not a parallel graph: one route at a time, one writing run per checkout. The only parallel work is up to four read-only readers in an analysis run, an investigation you ask for. With two or more routes, the host and the Architect check the combined change once more; a failure there stops the run.

The host stops every run at 500,000 tokens, 60 minutes of active time or 40 attempts.

Our simulator runs 36 scripted changes in seven languages through the real gate. We reported on 2026-09-24 that the gate got all 36 right with Jev and 34 on code alone, missing the Maven and Gradle build-file changes. We wrote those cases; we haven’t yet measured the gate on real projects, or Jev against its bars.

Try it

Signed in to ChatGPT and Anthropic, with GPT-6 Luna as your chat model, ask for a change that names a check in backticks, say pnpm test. If the plan skips it, the host sends it back once before any worker starts, Jev or not.

The working

Each number is from muxcode 0.12.7’s code, read on 2026-09-28, unless noted.

  • 0.7, 0.8, 0.85: Jev’s probability bars; policy, not yet measured.
  • 4 readings at medium: sure at 0.92 to 0.98, in the first live run of the shipped worker.
  • 500,000 tokens, 60 minutes, 40 attempts: cache reads and pauses don’t count; every step is an attempt, the gate included.
  • 36 and 34 of 36: from the changelog of 2026-09-24, with live Jev and without; not re-run here.