Hermes "⚠️ No response from provider for 180s (model: ..., context: ~... tokens). Reconnecting..."
Last checked
The error
⚠️ No response from provider for 180s (model: <model>, context: ~<N> tokens). Reconnecting...Status line from the stale-stream detector. The non-streaming variant reads "⚠️ No response from provider for 90s (non-streaming, model: <model>). Aborting call." After 5 stale kills in a row Hermes stops with "Provider has been unresponsive (no response received) for 5 consecutive stale attempts".
Hermes opened a streaming request and received no content for 180 seconds, so its stale-stream detector killed the connection and reconnected. 180s is the default for hosted endpoints with contexts under 50K tokens. If the provider is just slow, raise the limit with providers.<id>.stale_timeout_seconds in config.yaml or HERMES_STREAM_STALE_TIMEOUT in ~/.hermes/.env. If it is dead, fix the endpoint or switch models.
Why it happens
The stale-stream detector kills a stream that gets keep-alive pings or nothing at all but no actual tokens within the deadline. The default is 180s, raised to 240s above 50K tokens and 300s above 100K, with longer floors for known reasoning models.
- A hosted provider accepted the request but queued it or stalled, so no tokens arrived for three minutes. Busy shared or free routes are the usual suspects.
- Your model runs on a machine Hermes does not treat as local, such as a public hostname, a tunnel URL or a VPS. Loopback, private LAN, .local and Tailscale addresses get a 900s ceiling instead of 180s, but a public address gets the cloud default.
- The endpoint is dead or wedged: the server crashed after accepting the connection, or a proxy holds the socket open without forwarding anything.
- A slow thinking model on a large prompt took longer than the deadline to send its first token.
The fix
- 1 Check the log line to see which model and context size were involved: grep -i 'stale' ~/.hermes/logs/agent.log ~/.hermes/logs/gateway.log.
- 2 Confirm the endpoint is alive with a direct curl request to it. If it does not answer, restart the server or switch models with /model.
- 3 If the provider is just slow, raise the deadline for that provider in ~/.hermes/config.yaml: providers: <id>: stale_timeout_seconds: 600 (or per model under providers.<id>.models.<model>.stale_timeout_seconds).
- 4 Or raise it globally: add HERMES_STREAM_STALE_TIMEOUT=600 to ~/.hermes/.env. This only changes the 180s base; an explicit provider value always wins.
- 5 For a self-hosted model on a public URL, prefer the per-provider setting so you do not slow failure detection for every other provider.
- 6 Restart Hermes or run hermes gateway restart so the new values load.
echo 'HERMES_STREAM_STALE_TIMEOUT=600' >> ~/.hermes/.envWhy fallback does not kick in right away
A stale kill is treated as a dropped connection, so Hermes reconnects to the same provider first. Cross-provider fallback (fallback_providers in config.yaml) engages after the retry budget, agent.api_max_retries (default 3), is spent. With the defaults that can mean several 180s waits before you see a switch.
To fail over faster, set agent.api_max_retries to 0 so the first transient failure on the primary hands off to your fallback. Separately, HERMES_STREAM_STALE_GIVEUP (default 5) aborts calls immediately once that many stale kills happen in a row with no completed response.
Issue #7230, which is often cited for this, is about fallback not firing on auth errors at credential resolution. It is closed and has nothing to do with timeouts.
Still failing?
- If the message shows 900s instead of 180s, Hermes already treats your endpoint as local, and the server is genuinely stuck rather than slow.
- If tokens normally stream but stop mid-answer, look for a proxy or load balancer with an idle timeout shorter than the model's thinking time.
- If every provider times out, check the machine's network and DNS, since the problem is not the provider.
Related errors
Hit a different error?
Paste any agent error and get the cause and fix in seconds.
Frequently asked questions
Does this fire for local Ollama?
Rarely at 180s. For loopback, private LAN and Tailscale addresses Hermes raises the stale-stream ceiling to 900s (agent.local_stream_stale_timeout or HERMES_LOCAL_STREAM_STALE_TIMEOUT). If you set HERMES_STREAM_STALE_TIMEOUT yourself, that value replaces the local 900s too.
Is 180 seconds hard-coded?
No. It is the default for HERMES_STREAM_STALE_TIMEOUT. An explicit providers.
What is the 90s non-streaming version?
The non-streaming stale detector, which defaults to 90s (HERMES_API_CALL_STALE_TIMEOUT or the same stale_timeout_seconds key). It is disabled on local endpoints unless you set it explicitly.
Stop firefighting agent errors
Decoding errors one at a time is the manual version of what BetterClaw automates. Run your agents on a no-code AI agent platform with managed models, retries and config validation built in.
Free plan available · Pro $49/mo · BYOK · 7-day money-back guarantee
