Effort, one generation at a time
How muxcode picks a seat's reasoning effort before every model call, with the numbers from its first live runs.
In a muxcode workflow, Jev, the classifier, can move a seat’s reasoning effort before every model call, not once per task. In the two live runs whose readings we kept, it moved no level.
A seat is a role with its own model. The Architect plans, reviews its own plan, replans and judges. The worker writes the change. The retry takes over after a rejection or failed check. A roster is a workflow’s set of seats; muxcode ships a default. A generation is one model call. Reasoning effort is how hard the model thinks: low, medium, high, xhigh or max.
pi, the open-source coding agent muxcode runs, does the work. Jev is TypeSafe’s System One model: it answers typed questions with a probability, or a choice and its confidence. The host, muxcode’s own code around pi, sets the bar an answer must reach and applies the answer.
The checkpoint
Before each generation, once any tool results are in, Jev reads a bounded view of the seat’s work, never its private reasoning:
- the seat’s prompt, including any contract (the Architect’s written terms for the change), up to 12,000 characters;
- new input, and the seat’s public replies;
- the last 6 tool calls, each result cut to 4,000 characters from its start and end;
- a count of new tool failures, and which context window the seat is on.
A context window is everything the model reads on one request. The host trims the oldest entries to fit 60,000 bytes and redacts secrets.
Two questions, one bar
Jev answers two separate questions: which level the next generation needs, and the lease. A lease is how many generations that level should hold: 1, 2, 5 or 10, counting the next.
Each answer counts only at a confidence of 0.85 or more, out of 1. Below that, the level stays put and the lease is 1 generation.
During a lease the checkpoint takes no reading. New input, a failed tool call, a model change or a level set elsewhere ends it early. So does a fresh context window, which pi opens when one fills, starting from a handoff, a short record of the task.
Floors and ceilings
Jev moves a seat only inside a band the host sets. A judge never drops below its starting level, or high if that’s unknown: nothing after the judge catches an accept it got wrong. A worker never drops below its pin, the level the roster starts it at. A low pin saves effort, and the checks, the judge and the retry catch an attempt that reasoned too little.
A ceiling, “Up to” in Settings, caps the band and makes the pin the floor on every step. The host offers Jev only the levels between, and never asks about a one-level band.
When Jev is slow or silent
A missing reading leaves the level where it is. Each effort reading waits up to 2.5 seconds per attempt, with one retry. After two misses in a row, the host skips that seat’s next 5 checkpoints, then asks again. Each pause that ends in another miss doubles, up to 40 checkpoints.
In QA on 2026-09-23, 17 of about 65 live Jev calls missed the 2.5-second wait, 12 of them effort readings. The next day we gave the once-per-run questions 8 seconds; effort readings kept 2.5, since each holds a model call back. We don’t know yet how often they miss now.
The idea, and the dials
We based the checkpoint on Ares, by Jingbo Yang and colleagues. Its lightweight router predicts the lowest reasoning level each agent step needs, from the history so far. The authors report up to 52.7% fewer reasoning tokens than fixed high effort, with minimal loss in task success.
In other tools, a person chooses. Amp’s dial, since 2026-07-09, offers low, medium, high and ultra. Claude Code’s opusplan plans on Opus and executes on Sonnet. Claude’s adaptive reasoning varies thinking per step, but within the effort level a person set.
What the shipped roster does today
On the shipped roster, Jev reads two seats and never takes either below where it starts. The Architect, Claude Opus 5.5, is fixed at high, a one-level band, so Jev never reads it. The worker, also Opus 5.5, runs medium up to xhigh. The retry, Claude Fable 5.1, runs xhigh up to max.
Here Jev can only add effort above the pin, the reverse of Ares, which cuts it. Only the medium pin we set by hand saves anything against fixed high. Set the worker to low, and Jev raises its hard stretches from there.
In a QA run in mux’s cloud sandbox on 2026-09-23, on a test roster, Jev gave 8 level answers at 0.23 to 0.79 confidence. None cleared 0.85, so no level moved, and every lease came back as 1 generation.
The first live run with the shipped Architect and worker, filed 2026-09-24, had 4 worker readings. Each chose medium, its current level, at 0.92 to 0.98 confidence, in 208 to 252 ms. Every lease was 1 generation, at 0.36 to 0.61 confidence. Neither run reached the retry, so its band has no live reading yet.
We haven’t measured what this saves yet.
What you can change
You set each seat’s effort and “Up to” in Settings → Agents. Two changes widen what Jev reads:
- Raise the Architect’s “Up to” above high, and Jev reads its steps too, never below high.
- Add a cheaper worker ahead of the last in the worker list. For a contract the Architect marked routine and backed with notes from the code, Jev then picks a worker when it’s at least 0.85 sure. A high-risk review, or an unsure or silent Jev, leaves it with the last.
Jev runs for projects on a mux Pro or Team plan, or anywhere with your own TypeSafe key, set in Settings → Classifier. Without either, every seat keeps the effort you set.
Try it
After a few workflow runs, open Settings → Classifier and see where Jev’s Step effort answers fall against the 0.85 bar.
The working
Each number is set in muxcode 0.12.7’s code, published 2026-09-27 and read on 2026-09-28, unless noted.
- 17 of about 65 calls: timeouts in the 2026-09-23 QA report; we didn’t keep its log.
- The test roster: GLM 5.3 planned and judged, and DeepSeek V4 Flash worked.
- The retry’s band: moved up one level on 2026-09-24, after the live run.
- 52.7%: the Ares paper’s best case, submitted 2026-03-09.