Ollama

Ollama "Error: 500 Internal Server Error: memory layout cannot be allocated"

Last checked

The error

Error: 500 Internal Server Error: memory layout cannot be allocated

memory layout cannot be allocated with num_gpu = 26

Printed by ollama run or returned to your agent as an HTTP 500. The second form appears when num_gpu was set explicitly.

Ollama planned how to split the model between GPU and system memory, tried to allocate that plan, failed, retried with smaller layouts, and gave up. The message only exists in Ollama 0.24 and earlier: Ollama 0.30.0 (13 May 2026) moved model loading to a llama.cpp based runner and the string is gone from current builds. Upgrade to the current release (0.40.2 on 9 Oct 2026), or lower the context length and drop any fixed num_gpu.

Why it happens

Before 0.30, Ollama's own engine computed a GPU/CPU layer layout, allocated it, and backed off if the allocation failed. When the allocation still failed after the backoff, or failed at all with num_gpu fixed, it returned this error as a 500.

  1. A context window too large for the card. The KV cache is sized from the context length, and the VRAM based default can be 256K on large GPUs. The reporter of issue #14980 (Linux, 0.18.2) found the context was too large and their environment variables were not being read, and a maintainer suggested dropping OLLAMA_CONTEXT_LENGTH to 32768 on issue #15517.
  2. An explicit num_gpu. When num_gpu is set, Ollama skips the backoff and fails immediately with "memory layout cannot be allocated with num_gpu = N" if those N layers do not fit.
  3. An allocator bug on Windows in the 0.19 to 0.20 builds. Issue #15352 (0.20.2, RTX 3090, still open) shows small allocations failing for every model over about 10 GB while plenty of VRAM was free; the reporter said 0.19.x had worked, though issue #15143 reports the same error on 0.19.0.
  4. Unified memory machines on Windows. Issue #15517 (AMD Ryzen AI Max+ 395, Ollama 0.20.5) hit it with the installer build; it was closed in July 2026 without a documented code fix.
  5. Other processes holding VRAM, so the free memory Ollama measured at layout time was gone by allocation time.

The fix

  1. 1 Check your version with ollama -v. If it is below 0.30.0, upgrade to the current release from ollama.com or the GitHub releases page. This is the real fix: the code path that prints this error no longer exists.
  2. 2 If you must stay on an old build, lower the context length on the server, for example OLLAMA_CONTEXT_LENGTH=8192, then restart Ollama and reload the model.
  3. 3 Remove num_gpu from your Modelfile, API options or agent config so Ollama can back off on its own. If you need it, lower the value until the model loads.
  4. 4 Close other GPU heavy apps (games, browsers with hardware acceleration, a second Ollama instance) and confirm free VRAM with nvidia-smi before loading.
  5. 5 Optionally reserve headroom with OLLAMA_GPU_OVERHEAD, which takes a byte count per GPU, so the planned layout is more conservative.
  6. 6 Try a smaller quantization of the same model if it still does not fit.
ollama -v

Is it fixed? Status as of 9 October 2026

The error string was present in llm/server.go through Ollama 0.24.0 and is absent from 0.30.0 onward, including 0.40.2, the latest release. Ollama 0.30 replaced the old loader with a llama.cpp based runner across Windows and Linux.

Issue #15352 itself is still open with no linked fix, because nobody closed it after the engine change. On 0.30 and later an equivalent memory failure looks different: the log says "llama-server reported out-of-memory during startup", and the same levers (context length, quantization, free VRAM) apply. Issue #17517 (Windows, 0.32.5) reports a large Qwen model loading worse after an update, so new builds are not free of memory problems, just of this message.

Still failing?

  • Run ollama ps after a load attempt and check the CONTEXT column: if it shows a huge window you did not ask for, your context setting is not reaching the server.
  • Set OLLAMA_DEBUG=1, reproduce once, and read server.log around the 500 for the allocation sizes that failed.
  • If only models above about 10 GB fail on Windows with 0.20.x, you are on the issue #15352 path and upgrading is the only clean way out.

Related errors

Full guideOllama's Default Context Window Silently Truncates Your Agent ConfigEvery error, one pageOllama & OpenClaw Hardware: RAM, GPU & Budget Guide by Model Size (2026)

Hit a different error?

Paste any agent error and get the cause and fix in seconds.

Open the decoder

Frequently asked questions

Should I roll back to 0.18.x?

Not as a first move. The error exists in 0.18.2 too (issue #14980) and dates back to at least October 2025 (issue #12580). Upgrading past 0.30 removes the code that produces it, which is a better bet than rolling back.

Does pinning num_gpu fix it?

Only if you pin it low enough. Setting num_gpu explicitly disables Ollama's automatic backoff, so a value that is too high fails straight away with the num_gpu = N version of this error.

My GPU shows plenty of free VRAM. Why does allocation fail?

On the Windows 0.20.x builds the allocator itself was failing small requests with VRAM free (issue #15352). Elsewhere the usual reason is the KV cache: a long context can need more memory than the weights.

Stop firefighting agent errors

Decoding errors one at a time is the manual version of what BetterClaw automates. Run your agents on a no-code AI agent platform with managed models, retries and config validation built in.

Free plan available · Pro $49/mo · BYOK · 7-day money-back guarantee