Pi: "Response was truncated before completion."
Last checked
The error
Response was truncated before completion.
Truncated response recovery failed after one compact-and-retry attempt.The first line appears in red under the assistant message in Pi's interactive mode. The second appears only if Pi's automatic compact-and-retry also ends on a length stop.
Pi shows this line when the provider ends a reply with a length stop: finish_reason "length" on OpenAI-compatible APIs, or the max-tokens stop on Anthropic, Google, Bedrock and Mistral. The model ran out of output room. Set a real maxTokens and contextWindow for the model in ~/.pi/agent/models.json (custom models default to 16,384 and 128,000), run /compact to free context, or lower /thinking.
Why it happens
Pi asks for the model's maxTokens on every request, but clamps it to the space left in the context window (context window minus the estimated conversation minus a 4,096-token safety margin). A full session or a low limit both shrink the room for the answer.
- The model's maxTokens is too low for the job. Models you add yourself in ~/.pi/agent/models.json get 16,384 output tokens and a 128,000-token context window if you leave those fields out, and a long file write can exceed 16K.
- The session is nearly full. Because Pi clamps the request to the remaining context, a long session leaves little output room even when maxTokens is large.
- Thinking eats the budget. On Anthropic models Pi fits the thinking budget inside the same maxTokens, capped at maxTokens minus 1,024, so a high /thinking level leaves less room for the visible answer.
- The provider caps output below what Pi expects. OpenRouter lists a per-model max completion size, and several free models cap it at 8,192 tokens. On Pi older than 0.75.5, Bedrock's 4,096-token default did the same (issue #4848).
The fix
- 1 Run /model to confirm which provider and model are active, then look up that model's real output limit at the provider.
- 2 For a model you defined, set both limits on its entry in ~/.pi/agent/models.json, for example contextWindow 64000 and maxTokens 32768. Opening /model reloads the file.
- 3 For a built-in model, add a modelOverrides block under the provider, keyed by the model id, with the corrected maxTokens. Do not set it higher than the provider allows, or the request fails instead.
- 4 Run /compact (optionally with instructions) or /new to start clean, so the clamp leaves more room for output.
- 5 Lower the thinking level with /thinking if a reasoning model spends the budget before answering.
- 6 Update Pi with pi update. Current releases compact and retry once automatically on early length stops (0.84.0) and label that recovery correctly (0.84.3).
pi updateHow Pi differs from Hermes finish_reason=length
Hermes and Pi react to the same provider signal in opposite ways. Hermes sends no output cap to OpenAI-compatible servers and, when a reply stops on length, asks the model to continue up to four times before printing "response remained truncated after N continuation attempts".
Pi always sends an output cap and never auto-continues. Since 0.84.0 it does one thing: if the reply stopped below the model's configured maxTokens, Pi treats it as context pressure, compacts the session and retries once (PR #7540). If the reply used the full maxTokens, there is no automatic retry and you just see the red line.
When the provider is the one truncating
If Pi believes a model has 16,384 output tokens but the upstream provider stops at 8,192, the reply stops below Pi's limit. Pi then compacts and retries, the provider stops at 8,192 again, and you get "Truncated response recovery failed after one compact-and-retry attempt." Compaction cannot help here.
Set maxTokens to the provider's real cap so Pi stops treating it as context pressure, or switch to a paid route or a model with a larger completion limit.
Making Pi finish the job
The partial reply stays in the session, so you can type continue and the model picks up from there. If the cut-off content was a tool call, Pi does not run the half-written call, so ask for the file in smaller parts or as a series of edits instead of one large write.
Still failing?
- Check that the provider id and model id in models.json match exactly, because an unknown modelOverrides id is ignored silently.
- For Ollama or another local server, raise the server's own context size, since Pi cannot change it through the OpenAI-compatible API.
- If only one provider truncates, test the same prompt on another route to confirm the cap is upstream rather than in Pi.
Related errors
Hit a different error?
Paste any agent error and get the cause and fix in seconds.
Frequently asked questions
Where is Pi's max output setting?
There is no global switch. Output is the maxTokens field on each model, set in ~/.pi/agent/models.json either on a models entry or through modelOverrides for built-in models. Custom models default to 16,384.
Does Pi continue automatically like Hermes?
No. Pi does at most one compact-and-retry, and only when the reply stopped below the model's maxTokens. Otherwise you ask it to continue yourself.
Is compaction.reserveTokens the output limit?
No. compaction.reserveTokens (default 16384) decides when automatic compaction triggers and sizes summaries. The per-request output cap is the model's maxTokens, clamped to the remaining context.
Stop firefighting agent errors
Decoding errors one at a time is the manual version of what BetterClaw automates. Run your agents on a no-code AI agent platform with managed models, retries and config validation built in.
Free plan available · Pro $49/mo · BYOK · 7-day money-back guarantee
