← Field notes

No compaction

When a task fills its context window, pi opens a fresh one with a short handoff instead of a summary; here is where that line sits, what crosses it and what it costs.

With fresh context windows on, the default, muxcode doesn’t compact: no model summarizes your conversation to free room. A context window is everything the model reads on one request. When one fills, pi, the open-source coding agent muxcode runs, opens a new one. It holds the system prompt, the tools and a handoff, a short record carrying the task over. The rest stays on your Mac, in notes and an append-only history the model reads back. We adapted this from Posthorse, an open-source pi extension.

Where the line sits

A window rolls over once it passes its size minus 16,384 tokens, which pi reserves for a summary. That’s pi’s own compaction line; muxcode doesn’t move it. A checkpoint reminder, a prompt to save notes and roll over, comes up to 32,000 tokens earlier.

Model Window (tokens) Rollover (tokens) Reminder (tokens)
Opus 5.5, Fable 5.1 1,000,000 983,617 (98%) 951,617
GPT-6 Luna 272,000 255,617 (94%) 230,056

pi also rolls over on a “too long” error, or when the model asks. The reminder comes at most once per window, never after a finished answer, and one huge turn can jump the line.

What the handoff carries

When pi rolls over on its own, the handoff records inputs, not progress:

  • the task, its previous window and where its history lives;
  • its notes’ names and revisions;
  • your first and latest requests, your latest answer and the latest input;
  • the last tool results, if no reply followed;
  • the previous handoff, marked possibly stale, or a pointer if automatic.

Older inputs fill leftover room. Each record stops at 4,000 characters, the whole handoff at 20,000 characters, or half the new window’s free room if less. The model’s own handoff, when it asks, has the same limit.

When the usual record won’t fit, a minimal handoff carries pointers to notes and history, plus as much of your latest request as fits.

On paid mux plans (Pro and Team) or with your own TypeSafe key, Jev, the classifier (TypeSafe’s System One model), answers muxcode’s questions with a probability. The handoff then carries up to six 600-character excerpts of tool results Jev reads as still needed. If Jev reads a turn as a clean stopping point, the reminder can come up to 32,000 tokens sooner.

Reading back

The model reads back through two tools. A notes tool keeps durable notes, every revision readable. A history tool searches this branch of the conversation, the whole task or up to 200 recent project sessions. It matches plain text, not meaning.

Turning windows off

With Fresh context windows off (Settings → Models), the task’s own model summarizes older conversation at the same line, keeping about the newest 20,000 tokens verbatim. Open tasks need an app restart. A task started with windows off has no notes or history tools to read back what the summary drops. While windows are on, muxcode refuses a manual compact, a summary you ask for.

What it costs

A fresh window makes no extra call to your model; a summary makes one large call, two if the cut splits a turn. Inside a window the prompt only grows at its end, so the provider’s prompt cache serves earlier tokens cheaply. At pi’s catalog API prices, that’s $0.20 per million tokens on Opus 5.5 and $0.25 on Fable 5.1, against $4 and $10 uncached. A fresh GPT-6 Luna window opened at an estimated 13,400 tokens in our QA on 2026-09-25.

A summary can’t reuse that cache: pi rewrites the conversation as text under its own prompt, uncached, with each tool result cut to 2,000 characters. On a full window that’s at most about 950,000 tokens: $3.80 on Opus 5.5 and $9.50 on Fable 5.1, before output. We haven’t measured a typical one.

Louis-François Bouchard found full history cheaper than summarizing while cached input costs under about $0.55 per million tokens, as both prices above do. At $0.50, the two roughly tied. On Gemini 3.5 Flash, full history also recalled more: 92% against 38% for a summarizing preset.

Reading back costs calls. At 5× compression, Shuyu Liu found GPT-5.5’s retrieval calls tripled, 21.0 to 63.9, while completion’s rise from 80% to 85% wasn’t significant. Fresh windows drop more, so we expect more reading back; we haven’t measured it.

No more bloat than compaction

A window grows exactly as far with fresh windows as with compaction.

pi’s built-in tools cap each result at 50 KB; an extension’s tools may not. Without a paid plan or your own key, muxcode takes nothing back out before the line. With either, muxcode trims two things on Jev’s reading:

  • stretches Jev reads as unrelated in a command or search result of 12,000 characters or more, keeping head and tail. History keeps none of the cut text; the model reruns a narrower command.
  • older results of 800 characters or more that Jev reads as no longer needed, from 919,617 tokens on a 1,000,000-token window. Each becomes a pointer into history, for one cache miss in all.

The case against

Several teams now build long threads around compaction. OpenAI trained GPT-5.1-Codex-Max for it. Cursor trained Composer to summarize itself, halving compaction errors with summaries a fifth as long. Amp swapped its handoff for automatic compaction at 90% full. At Anthropic, Prithvi Rajasekaran used resets because Sonnet 4.5 had “context anxiety”: it wrapped up early near its limit. Opus 4.5 largely stopped that, so he dropped resets for automatic compaction.

We don’t know yet whether windows beat summaries on quality.

On your next long task, ask the agent to “note the result, then start a new context” between phases.

The working

  • pi 0.87.1 defaults, unchanged in muxcode 0.12.7: 16,384 and 20,000 tokens, 2,000 characters, 50 KB. Windows and prices: pi’s catalog, unchecked with providers.
  • muxcode 0.12.7, read on 2026-09-28: the character limits, six excerpts, 200 sessions and every token threshold. Rollover is window minus reserve, plus one.
  • 13,400 tokens: estimated at four characters a token.
  • $3.80 and $9.50: a ceiling, 950,000 tokens (983,617 less the 20,000 kept and the prompt) at $4 and $10 per million tokens.
  • Dates: OpenAI Nov 19, 2025; Cursor Mar 17, Anthropic Mar 24, Amp May 6, Liu Aug 17 and Bouchard Aug 18, 2026.