Guides 18 min read

Ollama Fetch Failed, Connection Refused, "Could Not Connect": Every Fix (Updated Sept 2026)

"Could not connect to ollama app", fetch failed, ECONNREFUSED on 11434? Find your exact error and paste the fix. Docker, Windows, n8n, VS Code, OpenClaw.

Shabnam Katoch

Shabnam Katoch

Growth Head

Ollama Fetch Failed, Connection Refused, "Could Not Connect": Every Fix (Updated Sept 2026)

Your agent framework says it can't reach Ollama. The model was running five minutes ago. Now it's "connection refused." Below are 15 fixes, plus the Windows, browser and crash variants, each under the exact error message you're seeing, with a copy-paste fix for each.

It was 11 PM. The agent had been running fine all day. Qwen3.8 27B on Ollama, classifying emails, posting digests to Slack. Solid.

Then the Slack digest didn't arrive. I checked the agent logs. Error: fetch failed. Checked Ollama. curl localhost:11434 said connection refused. ollama ps said "could not connect to ollama app". The server had silently died because my laptop went to sleep for ten minutes and Ollama didn't recover on wake.

This is the most frustrating class of agent errors because everything looks right. The model is pulled. The config is correct. It was working an hour ago. And now... connection refused.

Here's every Ollama fetch failed and connection error variant, why it happens, and the copy-paste fix. Bookmark this page. You'll be back.

Find your error

Search this table for the message you're seeing (Ctrl+F works), then jump to its fix.

Error you seeCause, and where to fix it
could not connect to ollama app / ollama serverServer not running. Fix 10: start the app or the service
fetch failed / ECONNREFUSEDOllama not running. Fix 1: restart the app or service
connection refused from Docker, works in terminalOllama listens on 127.0.0.1 only. Fix 2: bind to 0.0.0.0
connection refused (Docker)App uses localhost inside a container. Fix 4
('172.17.0.1', 11434)Ollama listens on 127.0.0.1 only. Fix 12
Cannot connect to host ... ssl:defaultServer down, or the container can't reach the host. Fix 12
bind: address already in useOllama is already running. Fix 11: use it
port blocked by firewallFirewall on 11434. Fix 3
SSL/TLS errorHTTPS instead of HTTP. Fix 5
timeout on loadModel too large for memory. Fix 6
timeout after a few messagesContext window overflow. Fix 7
model not foundWrong tag or outdated Ollama. Fix 8
Hermes/OpenClaw can't connectWrong base URL. Hermes needs /v1, OpenClaw must not have it. Fix 9
Problem in node 'AI Agent': fetch failedn8n points at the wrong host. Fix 13
VS Code fetch failedExtension runs remotely or uses the wrong endpoint. Fix 14
Claude Code Connection refusedServer down or wrong base URL. Fix 15
target machine actively refused itWindows, server not running. Windows
NetworkError / Failed to fetch (browser)CORS. Browser apps
POST .../api/generate: EOFRunner crashed, usually out of memory. EOF
Engine protocol predict request failedThat's LM Studio, not Ollama. Not Ollama?
Verification failed: fetch failed (OpenClaw)OpenClaw can't reach Ollama during setup. OpenClaw

Find your error: lookup cards pairing each Ollama error message with the fix number that solves it

The 30-second diagnostic (run this first)

Before trying fifteen different fixes, run these three commands:

# 1. Is the server up? (the real test)
curl http://localhost:11434/api/version

# 2. Which models are downloaded?
ollama list

# 3. Which models are loaded in memory right now? (empty is fine)
ollama ps
  • If curl says "connection refused": the server isn't running or isn't listening where you think. Go to Fix 1.
  • If curl works in your terminal but not in your app: the app is looking at a different localhost. Go to Fix 4 (Docker) or the app sections (n8n, VS Code, Claude Code).
  • If ollama ps shows an empty table: nothing is wrong. It only lists models loaded in memory, and the model loads on the first request.
  • If ollama ps prints could not connect to ollama app (or ollama server on newer versions): that's a down server, not a model problem. Go to Fix 10.
  • If everything connects but it's slow or timing out: go to Fix 6 or Fix 7.

The 30-second diagnostic: curl the version endpoint to test the server, ollama list for downloaded models, ollama ps for loaded models, with an empty ollama ps table marked as normal

Fix 1: Ollama service isn't running (the #1 cause)

"fetch failed" / "ECONNREFUSED 127.0.0.1:11434"

Symptoms: curl localhost:11434 returns "connection refused." Your agent logs show "fetch failed" or "ECONNREFUSED." ollama ps prints "could not connect to ollama app."

Why it happens: Ollama crashed, was never started, or died after laptop sleep/restart. On macOS, the Ollama app might have quit. On Linux, the systemd service might have failed.

Copy-paste fix:

# macOS: Relaunch the Ollama app
# Or from terminal:
ollama serve &

# Linux: Restart the service
sudo systemctl restart ollama
sudo systemctl status ollama

# Verify it's running
curl http://localhost:11434/api/version

Make it permanent (Linux):

# Enable auto-start on boot
sudo systemctl enable ollama

On macOS, add Ollama to Login Items (System Settings → General → Login Items) so it starts automatically. On Windows, see the Windows section.

Fix 2: Ollama is running but bound to the wrong address

"connection refused" from Docker or another machine

Symptoms: curl localhost:11434 works on the machine itself, but your agent (running in Docker, WSL or on another machine) gets "connection refused."

Why it happens: By default, Ollama binds to 127.0.0.1:11434. This means only processes on the same machine can connect. If your agent runs in Docker, a VM, or a different machine, it can't reach 127.0.0.1.

Copy-paste fix:

# Linux: Set Ollama to listen on all interfaces
sudo systemctl edit ollama

# Add these lines:
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"

# Restart
sudo systemctl daemon-reload
sudo systemctl restart ollama

# macOS: Set the environment variable
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"
# Then quit and reopen the Ollama app

Security warning: 0.0.0.0 exposes Ollama's API, which has no authentication, to your whole network, and to the internet on a VPS with an open port. Only do this behind a firewall, or bind to the default Docker bridge IP (usually 172.17.0.1:11434) instead so only containers can reach it.

Verify from inside the container, because testing from the host proves nothing about Docker:

docker exec -it <container> curl http://host.docker.internal:11434/api/tags

If the container image has no curl, use wget instead:

docker exec -it <container> wget -qO- http://host.docker.internal:11434/api/tags

The single most common reason agents in Docker can't reach Ollama: the server is bound to 127.0.0.1 (localhost only). Agents in Docker need Ollama on 0.0.0.0 or the bridge IP.

Fix 3: Port 11434 is blocked or in use

Symptoms: Ollama starts but other machines can't reach it, or another process is sitting on port 11434.

Copy-paste fix:

# Check what's using port 11434
lsof -i :11434   # macOS/Linux
ss -tlnp | grep 11434   # Linux

# If a process that is NOT Ollama holds it, move Ollama to another port
OLLAMA_HOST=127.0.0.1:11435 ollama serve

If the process on 11434 is Ollama itself, nothing is wrong. See Fix 11.

If a firewall is blocking the port:

# Linux (ufw)
sudo ufw allow 11434/tcp

# macOS: System Settings → Network → Firewall

Fixes 1 to 3, the server side: nothing is listening, listening on 127.0.0.1 only, or the port is blocked or held by another app, each with its copy-paste command

Fix 4: Agent is using the wrong URL (Docker networking)

"fetch failed" in the app, but ollama run works

Symptoms: Ollama works fine in terminal (ollama run qwen3.8:27b works). But your agent framework (OpenClaw, Hermes, n8n) returns "fetch failed" or "connection refused."

Why it happens: Your agent is pointing to http://localhost:11434 but it's running inside Docker. Inside Docker, localhost means the container itself, not your host machine.

Copy-paste fix by framework:

OpenAI-compatible clients (Hermes and most agent frameworks) running in Docker:

http://host.docker.internal:11434/v1

host.docker.internal resolves to the host machine from inside a Docker container. Works on macOS and Windows Docker Desktop. On Linux, add it yourself:

# docker run
docker run --add-host=host.docker.internal:host-gateway ...
# docker compose
extra_hosts:
  - "host.docker.internal:host-gateway"

On Linux, Ollama on the host also has to listen beyond 127.0.0.1, or the container reaches the host and still gets refused (Fix 2). Docker Desktop on macOS and Windows forwards host.docker.internal to the host's localhost, so the default bind works there.

n8n (running in Docker): see Fix 13.

Any framework (running natively, not Docker): use http://localhost:11434 or http://127.0.0.1:11434. Don't put 0.0.0.0 in a client URL. It's a bind address for servers, not an address to connect to.

Where localhost points: on the host it reaches Ollama directly, inside a Docker container it needs host.docker.internal, and in WSL2 or a VS Code remote it needs mirrored networking or Ollama on the same side

Fix 5: HTTPS vs HTTP mismatch

Symptoms: "fetch failed" with an SSL/TLS error in the logs. Or the agent framework requires HTTPS but Ollama serves HTTP.

Why it happens: Ollama serves HTTP by default (no SSL). Some agent frameworks or reverse proxies expect HTTPS.

Copy-paste fix: Use http:// not https:// in your Ollama URL. If you need HTTPS (for remote access), put a reverse proxy (Nginx, Caddy) in front of Ollama:

# Caddy (simplest HTTPS reverse proxy)
# In Caddyfile:
ollama.yourdomain.com {
    reverse_proxy localhost:11434
}

A Python error that says ssl:default is not this problem. See Fix 12.

Fix 6: Model too large for available memory (silent OOM)

Symptoms: Ollama starts. The model begins loading. Then... silence. No response. Eventually: timeout. Or the model loads but inference is impossibly slow (under 1 tok/s). Or the connection drops with EOF.

Why it happens: The model's memory requirements exceed your available RAM or VRAM. Ollama doesn't always error clearly. It just hangs, spills onto the CPU, or the runner dies.

Copy-paste fix:

# Check what's loaded, its size, and whether it spilled to CPU
ollama ps

# If the model is too large, use a smaller quant
ollama pull qwen3.8:27b-q4_K_M    # about 18 GB
# Instead of:
# ollama pull qwen3.8:27b-q8_0    # about 30 GB

A mixture-of-experts model does not save memory here. qwen3.6:35b-a3b only activates 3B parameters per token, but all 23 GB of weights still have to fit. See our local LLM hardware guide for what fits at each RAM tier.

Fix 7: Context window overflow causes timeout

Symptoms: First few messages work fine. Then after 5-10 messages, responses stop or time out. No error. Just... waiting.

Why it happens: The conversation grew beyond the model's allocated context window. Ollama tries to process all tokens and runs out of memory or slows to a crawl.

Copy-paste fix:

# Set a context window in a Modelfile
# Don't use the maximum, use what you need

cat > agent.modelfile << 'EOF'
FROM qwen3.8:27b
PARAMETER num_ctx 32768
PARAMETER num_predict 2048
EOF

ollama create agent -f agent.modelfile

Then point your agent at the new agent model instead of the original tag, or the setting won't apply.

32768 is the floor for simple agents. Hermes Agent needs at least 64K for tool use, so set PARAMETER num_ctx 64000 for Hermes, and give Claude Code at least as much. Every extra token of context costs memory, so if raising num_ctx tips you into Fix 6, step down a model size.

Our context window management guide covers how to prevent context overflow, and our guide to Ollama's default context window explains why the defaults are too small for agent tasks.

Fixes 6 and 7: the model does not fit in memory, or the conversation outgrows num_ctx, with the commands and context sizes for each

Fix 8: Ollama version mismatch (model not found)

"Error: pull model manifest: file does not exist" / "model not found"

Symptoms: ollama run modelname returns "model not found" even though you pulled it. Or the model exists in ollama list but won't load. Or a brand new model refuses to pull.

Why it happens: New models often use new architectures, and an older Ollama can't load them. The tag in your agent config also has to match ollama list exactly.

Copy-paste fix:

# Check your version
ollama --version

# Update Ollama
# macOS / Windows: the app auto-updates, or reinstall from ollama.com
# Linux:
curl -fsSL https://ollama.com/install.sh | sh

# Re-pull the model after updating
ollama pull qwen3.8:27b

If a pull was interrupted, the partial download can leave a broken model behind. Remove it and pull again:

ollama rm qwen3.8:27b
ollama pull qwen3.8:27b

Fix 9: Hermes/OpenClaw specific connection issues

Symptoms: Hermes or OpenClaw can't connect to Ollama, even though curl and ollama run work fine from the terminal.

Why it happens: Each framework wants the base URL in a specific format, and the two don't agree on /v1.

Hermes Agent: "Connection error" or "fetch failed"

Hermes uses Ollama's OpenAI-compatible endpoint, so its base URL ends in /v1. Hermes is configured in ~/.hermes/config.yaml:

model:
  default: "qwen3.8:27b"
  provider: "custom"
  base_url: "http://localhost:11434/v1"

Leave the API key empty or set it to no-key. If Hermes runs in Docker, swap localhost for host.docker.internal and keep the /v1. Hermes needs at least 64K of context for tool use (Fix 7).

If tool calls hang with Qwen3.8 on /v1, update Ollama first, then see our Qwen3.8 27B tool calling fix. qwen3.6:27b is the fallback that works on the same endpoint. For other Hermes errors, see our error 400 diagnostic guide, and if Ollama connects but local models hang on tool calls, our guide to Hermes Agent not working covers the failures specific to local backends.

OpenClaw: "TypeError: fetch failed"

OpenClaw is the opposite. Its Ollama provider talks to Ollama's native API, so the base URL has no /v1. OpenClaw's docs warn that /v1 breaks tool calling:

{
  "models": {
    "providers": {
      "ollama": {
        "baseUrl": "http://127.0.0.1:11434",
        "api": "ollama",
        "apiKey": "ollama-local"
      }
    }
  }
}

The easiest route is ollama launch openclaw --model qwen3.8:27b, which writes this config for you. For every OpenClaw variant (model discovery timeouts, TUI errors, WSL2), see our OpenClaw Ollama fetch failed guide, and for tool-calling and "no response" problems, our OpenClaw local model not working guide.

The switchboard: which Ollama URL each client needs, with /v1 for OpenAI-compatible clients like Hermes, no /v1 for native clients like n8n, OpenClaw and Claude Code, and host.docker.internal inside Docker

Fix 10: "could not connect to ollama app"

"Error: could not connect to ollama app, is it running?"

Newer Ollama versions word it as Error: could not connect to ollama server, run 'ollama serve' to start it. Same meaning: the Ollama CLI is telling you the server isn't up. The CLI (ollama run, ollama ps, ollama pull) is only a client. It needs the background server on port 11434.

Windows: check the system tray for the llama icon. If it's missing, open Ollama from the Start menu. Still failing? Quit it from the tray, open a new terminal and run ollama serve to see the actual error.

macOS: open the Ollama app from Applications. The menu bar icon means the server is running. (On macOS and Windows, ollama ps and ollama list try to start the app for you, so if you still see this error, the app is failing to start. Check the log in the EOF section.)

Linux:

sudo systemctl start ollama
sudo systemctl enable ollama
journalctl -u ollama -n 50 --no-pager   # see why it died

If you set OLLAMA_HOST to a custom port, the CLI needs the same variable, or it'll look on 11434 and fail.

Fix 11: "bind: address already in use" on 11434

"Error: listen tcp 127.0.0.1:11434: bind: address already in use"

This almost always means Ollama is already running. The desktop app or the systemd service started it, and you ran ollama serve a second time.

You don't need to fix anything. Just use it:

curl http://localhost:11434/api/version

If you really want to run ollama serve by hand (for example, to set env vars), stop the other one first:

# Linux
sudo systemctl stop ollama
# macOS / Windows: quit Ollama from the menu bar or system tray

Only if lsof -i :11434 (macOS/Linux) or netstat -ano | findstr 11434 (Windows) shows a process that isn't Ollama, move Ollama to another port (Fix 3).

Fix 12: Python "Cannot connect to host localhost:11434 ssl:default"

"Cannot connect to host localhost:11434 ssl:default [Connect call failed ('127.0.0.1', 11434)]"

This is aiohttp (used by Open WebUI, LiteLLM and many async Python agents) saying the TCP connection was refused. The ssl:default part is just aiohttp's label. It does not mean an SSL problem.

  • ('127.0.0.1', 11434): the server isn't running (Fix 1 / Fix 10), or your Python code is in a container where localhost is the container.
  • host.docker.internal:11434 ... ('172.17.0.1', 11434): your container found the host fine, but Ollama on the host is only listening on 127.0.0.1. Set OLLAMA_HOST=0.0.0.0:11434 (Fix 2) and restart Ollama.
  • host.docker.internal ... [Domain name not found]: Linux Docker doesn't create that name by default. Add --add-host=host.docker.internal:host-gateway or, in compose:
extra_hosts:
  - "host.docker.internal:host-gateway"

Fix 13: n8n "fetch failed" / "ECONNREFUSED"

"Problem in node 'AI Agent': fetch failed" / "Couldn't connect with these settings ECONNREFUSED"

The n8n Ollama credential test runs from where n8n runs, not from your browser.

Set the Base URL in the Ollama credential to match your setup:

# n8n and Ollama both installed natively
http://localhost:11434

# n8n in Docker, Ollama on the host
http://host.docker.internal:11434

# Both in the same docker compose (use the service name)
http://ollama:11434

On Linux, the Docker case also needs the extra_hosts line from Fix 12 and Ollama bound beyond 127.0.0.1 (Fix 2). No /v1 here. The n8n Ollama node uses Ollama's native API. If the credential saves but the AI Agent node fails mid-run, it's usually a timeout on a slow model. Use a smaller model or raise the n8n execution timeout. Our n8n vs Make comparison covers the full agent setup.

Fix 14: VS Code "Ollama fetch failed"

"fetch failed" in a VS Code extension

VS Code extensions (Continue, Copilot's bring-your-own-model, Cline and others) connect from the machine the extension runs on.

  • Remote SSH, WSL or Dev Containers: localhost now means the remote machine or container, not your laptop. Either run Ollama on that remote, or point the extension at your host's IP with Ollama bound to 0.0.0.0 (Fix 2, and read its security warning).
  • Check the endpoint field. Most want http://localhost:11434 (no /v1). Extensions that use the OpenAI format want http://localhost:11434/v1.
  • Test from VS Code's own terminal. If the command below fails there, it's a network location issue, not the extension.
curl http://localhost:11434/api/version

Fix 15: Claude Code "Connection refused"

"Connection refused, a firewall or proxy may be blocking it"

If you're running Claude Code on Ollama models, the easiest path is Ollama's own launcher, which sets everything up:

ollama launch claude --model qwen3.8:27b

Setting it up by hand? Claude Code speaks the Anthropic API, which Ollama serves natively, so the base URL takes no /v1:

export ANTHROPIC_BASE_URL=http://localhost:11434
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""

Connection refused here is the same root cause as everywhere else: the server isn't up, or Claude Code runs somewhere (container, WSL) where localhost isn't your Ollama. Give it at least 64K of context for real repos (Fix 7).

Windows: "target machine actively refused it"

"No connection could be made because the target machine actively refused it"

That's Windows for "connection refused". Checks:

  1. Is Ollama in the system tray? If not, start it (Fix 10).
  2. Test with curl.exe, not curl. In PowerShell, curl alone is an alias for Invoke-WebRequest. Run curl.exe against 127.0.0.1:11434/api/version, as shown below.
  3. To set OLLAMA_HOST on Windows: quit Ollama from the system tray first. Open Settings (Windows 11) or Control Panel (Windows 10), search for "environment variables", choose "Edit environment variables for your account", add OLLAMA_HOST, then start Ollama again from the Start menu.
  4. WSL2: an app inside WSL can't reach Windows Ollama on localhost by default. Either install Ollama inside WSL, or turn on mirrored networking (networkingMode=mirrored under [wsl2] in .wslconfig) and run wsl --shutdown to restart WSL. Mirrored mode needs Windows 11 22H2 or later.
curl.exe http://127.0.0.1:11434/api/version

For firewall, antivirus and PowerShell-specific fixes, see our Ollama connection refused on Windows guide.

Browser apps and extensions: "NetworkError" / "Failed to fetch"

"NetworkError when attempting to fetch resource" / "TypeError: Failed to fetch"

Ollama is running, but it's rejecting requests from a browser origin it doesn't trust (CORS). Pages on localhost and 127.0.0.1 are allowed by default. Browser extensions and web apps on any other domain are not, so allow them explicitly:

# Linux (sudo systemctl edit ollama)
Environment="OLLAMA_ORIGINS=chrome-extension://*,moz-extension://*,safari-web-extension://*,https://your-app.com"

# macOS
launchctl setenv OLLAMA_ORIGINS "chrome-extension://*,moz-extension://*,safari-web-extension://*"

Restart Ollama after. On Windows, add OLLAMA_ORIGINS as a User environment variable the same way as OLLAMA_HOST above.

Runner crashed: "POST .../api/generate: EOF" or "server disconnected"

"Error: POST .../api/generate: EOF" / "Server disconnected without sending a response"

The connection opened, then the model runner died mid-response. Nine times out of ten it's memory: the model plus its context doesn't fit. See Fix 6 and Fix 7. Check the log for "out of memory":

# Linux
journalctl -u ollama --no-pager | grep -i "out of memory"
# macOS
grep -i "out of memory" ~/.ollama/logs/server.log
# Windows (PowerShell)
Select-String -Path "$env:LOCALAPPDATA\Ollama\server.log" -Pattern "out of memory"

Not actually Ollama?

OpenClaw: "Verification failed: fetch failed"

"Verification failed: fetch failed" / "model verification failed"

OpenClaw checks the model during setup by calling your endpoint. "Fetch failed" at this step means it couldn't reach Ollama at all, not that the model is bad (people often search it as "qwen3.8-27b model verification failed").

  1. Confirm curl http://127.0.0.1:11434/api/tags lists the exact model tag you typed. qwen3.8:27b (colon) is not qwen3.8-27b (hyphen).
  2. Make sure the base URL has no /v1 (Fix 9).
  3. Or skip the manual setup: ollama launch openclaw --model qwen3.8:27b lets Ollama write the config for you.

For every other OpenClaw variant, see our full OpenClaw Ollama fetch failed guide.

The honest truth about local agent backends

The honest truth about running local models as agent backends: connection errors are part of the deal. Ollama is excellent software. But it runs on your machine, depends on your network configuration, and interacts with your Docker setup, your firewall, your sleep settings, and your memory limits. Every one of these is a potential failure point.

For agents that need to run reliably 24/7, the question isn't "can I fix this error?" It's "do I want to keep fixing these errors?"

Give BetterClaw a look if you'd rather build agents than debug connection strings. Free plan with 1 agent and 100 credits a month. Pro is $49/month for 5 agents. 200+ verified skills. 28+ model providers via BYOK with zero markup. We handle the connections. You handle the agent logic.

Frequently Asked Questions

Why does Ollama say "fetch failed"?

"Fetch failed" means the client (your agent framework, browser, or CLI tool) couldn't establish a connection to Ollama's HTTP server. The three most common causes: the Ollama service isn't running (Fix 1), the server is bound to localhost but the client is in Docker (Fix 4), or the port is blocked by a firewall (Fix 3). Run curl http://localhost:11434/api/version to check if the server is responding.

What does "could not connect to ollama app, is it running?" mean?

The Ollama CLI can't reach the background server on port 11434. Newer versions say "could not connect to ollama server" instead. Start the Ollama app (Windows, macOS) or run sudo systemctl start ollama (Linux), then retry.

Why does "ollama ps" show nothing?

ollama ps only lists models loaded in memory. An empty list is normal when nothing has been used recently. Models load on the first request. To check the server itself, run curl http://localhost:11434/api/version.

What does "bind: address already in use" mean in Ollama?

Ollama is already running, usually started by the desktop app or systemd. You don't need a second ollama serve. Use the running one, or stop it first.

How do I fix Ollama connection refused in Docker?

Docker containers can't reach localhost on the host machine. Replace http://localhost:11434 with http://host.docker.internal:11434 in your agent framework's config. On Linux Docker, add --add-host=host.docker.internal:host-gateway to your Docker run command, and make Ollama accept connections from the container by setting OLLAMA_HOST=0.0.0.0:11434 behind a firewall. Docker Desktop on macOS and Windows works with the default bind.

Does Ollama need /v1 at the end of the URL?

Only for clients that use the OpenAI-compatible format, such as Hermes Agent and most agent frameworks: http://localhost:11434/v1. Clients that use Ollama's native API take the base URL with no /v1: OpenClaw's Ollama provider, the n8n Ollama node, Claude Code via ANTHROPIC_BASE_URL, and the Ollama Python library. Adding /v1 to OpenClaw breaks tool calling.

Why does Ollama time out after a few messages?

Conversation context grows with each message. By message 10-15, the accumulated tokens may exceed Ollama's allocated context window, causing it to hang or become extremely slow. Fix: set PARAMETER num_ctx 32768 in your Modelfile as a floor, or 64000 for Hermes Agent. Also set PARAMETER num_predict 2048 to cap output length. If the model is too large for your RAM, switch to a smaller model or quantization.

How do I connect Hermes or OpenClaw to Ollama?

For Hermes: in ~/.hermes/config.yaml, set provider: "custom" and base_url: "http://localhost:11434/v1" under model, and leave the API key empty. For OpenClaw: set baseUrl to http://127.0.0.1:11434 with api: "ollama" and no /v1, or run ollama launch openclaw to write the config for you. If either runs in Docker, use host.docker.internal instead of localhost.

Should I use Ollama or a cloud API for agent backends?

Ollama is ideal for development, testing, privacy-sensitive workloads, and high-volume inference where API costs compound. Cloud APIs (via BYOK on platforms like BetterClaw) are better for production reliability (no connection errors, no sleep/wake issues, no port debugging), access to frontier models, and 24/7 uptime. Many teams use Ollama for development and cloud APIs for production.

Want to skip the setup?

BetterClaw does this in 60 seconds. No Docker, no config files.

Start free
Tags:ollama fetch failedollama connection refusedcould not connect to ollama appollama not respondingollama localhost 11434ollama timeoutollama dockern8n ollamahermes ollama connection
Share this article
Was this helpful?