Ollama

Ollama not using GPU (100% CPU)

Last checked

The error

NAME          ID              SIZE     PROCESSOR    CONTEXT    UNTIL
llama3:8b     365c0bd3c000    6.2 GB   100% CPU     4096       4 minutes from now

level=INFO source=types.go msg="inference compute" id=cpu library=cpu

ollama ps output plus the startup line in server.log when no GPU was accepted. Older builds (0.11 and earlier) logged "no compatible GPUs were discovered".

Ollama runs on the CPU when it either finds no GPU it can use at startup or the model does not fit in VRAM. Run ollama ps: 100% CPU means nothing went to the GPU, a CPU/GPU split means the model spilled. Then read the server log for the inference compute line. If it says id=cpu, fix the driver, the container GPU flags, or the unsupported card, and restart Ollama.

Why it happens

Ollama inventories GPUs once when the server starts and drops any device it cannot use, logging why. If every device is dropped, all inference runs on the CPU until the server restarts.

  1. NVIDIA driver too old or not loaded. Ollama needs compute capability 5.0+ and driver 550 or newer (570+ for compute 5.0 to 6.2) and logs "NVIDIA driver too old" with the required version. On Linux, suspend and resume can also lose the GPU until the nvidia_uvm module is reloaded.
  2. Docker without GPU access. The container needs the NVIDIA Container Toolkit and --gpus=all, or for AMD the rocm image with --device /dev/kfd --device /dev/dri. Docker Desktop on macOS has no GPU passthrough at all.
  3. AMD card or driver not supported. Ollama needs the ROCm v7 driver; it logs "AMD driver is too old" on Windows and drops cards whose gfx target has no rocBLAS support, suggesting HSA_OVERRIDE_GFX_VERSION. Linux users also need video/render group access to /dev/kfd.
  4. Integrated GPU skipped. Ollama drops some iGPUs by default and logs "dropping integrated GPU; to enable, set OLLAMA_IGPU_ENABLE=1".
  5. The model or context is bigger than VRAM, so layers spill to the CPU. ollama ps shows a split such as 48%/52% CPU/GPU rather than 100% CPU.

The fix

  1. 1 Load the model, then run ollama ps and read the PROCESSOR column: 100% GPU is fine, 100% CPU means no GPU was used, a split means not enough VRAM.
  2. 2 Open the server log (journalctl -u ollama on Linux, %LOCALAPPDATA%\Ollama\server.log on Windows, ~/.ollama/logs/server.log on macOS, docker logs for containers) and find the inference compute lines near startup.
  3. 3 NVIDIA: update the driver, confirm nvidia-smi works, and on Linux after sleep run sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm, then restart Ollama.
  4. 4 Docker: install the NVIDIA Container Toolkit, run sudo nvidia-ctk runtime configure --runtime=docker, restart Docker, and start the container with docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama. Test with docker run --gpus all ubuntu nvidia-smi.
  5. 5 AMD: install the ROCm v7 driver; for an unsupported gfx target set HSA_OVERRIDE_GFX_VERSION to the nearest supported one (for example 10.3.0) on the server.
  6. 6 If ollama ps shows a split, lower the context length or use a smaller quantization so the whole model fits in VRAM.
ollama ps

Log lines worth searching for

Search server.log for these strings, all present in current Ollama source: "inference compute" (one line per accepted device; id=cpu library=cpu means none), "NVIDIA driver too old", "AMD driver is too old. Update your AMD driver to enable GPU inference.", "no rocblas support for gfx target", "dropping integrated GPU", "user overrode visible devices", and "failure during llama-server GPU discovery".

A visible devices warning means CUDA_VISIBLE_DEVICES, ROCR_VISIBLE_DEVICES or a similar variable is set; an invalid value such as -1 deliberately forces CPU. Unset it and restart.

Still failing?

  • Set OLLAMA_DEBUG=1, restart the server, and read the extra GPU discovery output before the inference compute line.
  • On Windows, make sure the Ollama that answers on port 11434 is the one you configured: environment variables are read when the tray app starts, so quit it from the tray and start it again after changes.
  • In Docker on Linux, if the GPU works at first and then drops to CPU with discovery errors, the Ollama docs suggest setting the cgroupfs driver in /etc/docker/daemon.json.

Related errors

Full guideOllama & OpenClaw Hardware: RAM, GPU & Budget Guide by Model Size (2026)

Hit a different error?

Paste any agent error and get the cause and fix in seconds.

Open the decoder

Frequently asked questions

How do I know for sure whether Ollama is using my GPU?

Run ollama ps while a model is loaded. The PROCESSOR column says 100% GPU, 100% CPU, or a CPU/GPU split. Task Manager or nvidia-smi showing VRAM use by the Ollama process confirms it.

My GPU is detected but generation is still slow. Why?

The model plus its KV cache does not fit, so some layers run on the CPU. A long context window is the usual culprit. Lower it or pick a smaller quant until ollama ps shows 100% GPU.

Can I force Ollama onto the CPU on purpose?

Yes. The docs say to set CUDA_VISIBLE_DEVICES (or ROCR_VISIBLE_DEVICES for AMD) to an invalid ID such as -1. If you see 100% CPU unexpectedly, check that one of these is not set by accident.

Stop firefighting agent errors

Decoding errors one at a time is the manual version of what BetterClaw automates. Run your agents on a no-code AI agent platform with managed models, retries and config validation built in.

Free plan available · Pro $49/mo · BYOK · 7-day money-back guarantee