Skip to content

GGUF models on auto: start the runner again on the CPU when it dies on the GPU - #47

Merged
cobanov merged 1 commit into
mainfrom
gguf-cpu-retry
Oct 2, 2026
Merged

cobanov merged 1 commit into
mainfrom
gguf-cpu-retry

Conversation

@cobanov

@cobanov cobanov commented Oct 1, 2026

Copy link
Copy Markdown
Member

A GGUF runner on auto already falls back to the CPU itself when llama.cpp reports an error while loading or warming up on the GPU. It can't do that when the GPU backend takes the process down before the runner answers, for example a driver fault or an abort inside ggml (as with the CUDA kernels in #42). The server then reported a failed load, and every request for that model failed the same way.

The scheduler now handles GGUF models on auto the same way it handles ONNX models with a GPU pack. If the runner doesn't come up, the scheduler starts it again with --device cpu, using the same executable and environment, and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), because then any GPU driver on the machine is in the path.

Tested on Linux:

  • New HTTP test: the fake runner exits before answering unless it runs on the CPU. With this change the model loads on the CPU; without it the test fails.
  • A model that fails on every device still reports its own error.
  • clippy is clean and the ollaya-server tests pass.

…n the GPU

A GGUF runner on `auto` already falls back to the CPU itself when llama.cpp reports an error while
loading or warming up on the GPU. It cannot when the GPU backend takes the process down before
the runner answers: a driver fault, or an abort inside ggml, as with the CUDA kernels of #42. The
server then reported a failed load, and every request for the model failed the same way.

The scheduler now treats GGUF models on `auto` as it treats ONNX models with a GPU pack: if the
runner does not come up, it starts it again with `--device cpu` (same executable and environment)
and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), where
any GPU driver on the machine is in the path.

Tested: an HTTP test whose fake runner exits before answering unless it runs on the CPU loads on
the CPU (and fails without this change); a model that fails on every device still reports its
own error.
@cobanov
cobanov merged commit 50fdda2 into main Oct 2, 2026
6 checks passed
@cobanov
cobanov deleted the gguf-cpu-retry branch October 2, 2026 00:28
cobanov pushed a commit that referenced this pull request Oct 5, 2026
winnow:e4b-vision is the e4b tag (same GGUF, decision and calibration) plus the author's vision projector from the same pinned revision. PNGs given as images on /api/decide or --image on the command line, up to 16 per request, are read with the state and the questions. /v1/systemone keeps TypeSafe's text-only schema; the text tags keep their path.

The runner loads libmtmd from the pinned llama.cpp release next to libllama, prepares the images and the state once per request and evaluates each question from that split, as stock llama-server tokenizes a multimodal prompt. libmtmd's prompt-bearing debug messages stay out of the log. The projector stays on the runner's device, with the CPU fallback of #47.

Parity against stock llama-server b11146 with the projector on the same device: 21 requests and 65 questions, every decision the same, option logits within 9.6e-6 on an RTX 4070, 9.5e-6 on an RTX 5090 (the shipped 8,192-token context) and 1.2e-5 on an x86-64 CPU; the 505 text questions pass on both GPUs. docs/families/winnow.md has the numbers and the commands.
cobanov pushed a commit that referenced this pull request Oct 6, 2026
The Windows base install and desktop app take llama.cpp's win-vulkan-x64 build of the pinned b11146 instead of win-cpu: the same CPU backends plus ggml-vulkan.dll, so GGUF models use a GPU of any vendor on Windows. OLLAYA_DEVICE accepts vulkan and vulkan:<n>. Under auto, CUDA comes first when its pack is installed, then the discrete GPU with the most free memory; an integrated Vulkan GPU shares system memory and is not shown to beat the CPU, so it takes an explicit vulkan:<n> (CUDA's integrated devices, such as a DGX Spark's GB10, and Metal stay eligible). A GPU runner that fails while it loads starts again on the CPU (#47).

Parity is per device (ADR 0003, point 7): against stock llama-server b11146 on the same Vulkan device, an Intel Arc 140T gives winnow:e4b 505/505 within 1.1e-5 and jevk5 593/593 within 9.5e-6, and an RTX 4090 gives winnow:e4b 505/505, jeb:9b 494/494 and cygnet:12b 502/502, all within 1.1e-5. On a Windows build of this branch with an RTX 4090 and a UHD 770, auto picks the RTX 4090 through Vulkan, the CPU when only the UHD 770 is visible, and the UHD 770 with vulkan:0. docs/measurements/intel-arc-140t-parity.md records the Arc runs and their distance to CUDA, which is about the CPU's.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant