Repository navigation
GGUF models on auto: start the runner again on the CPU when it dies on the GPU - #47
Merged
Merged
Conversation
…n the GPU A GGUF runner on `auto` already falls back to the CPU itself when llama.cpp reports an error while loading or warming up on the GPU. It cannot when the GPU backend takes the process down before the runner answers: a driver fault, or an abort inside ggml, as with the CUDA kernels of #42. The server then reported a failed load, and every request for the model failed the same way. The scheduler now treats GGUF models on `auto` as it treats ONNX models with a GPU pack: if the runner does not come up, it starts it again with `--device cpu` (same executable and environment) and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), where any GPU driver on the machine is in the path. Tested: an HTTP test whose fake runner exits before answering unless it runs on the CPU loads on the CPU (and fails without this change); a model that fails on every device still reports its own error.
cobanov
pushed a commit
that referenced
this pull request
Oct 5, 2026
winnow:e4b-vision is the e4b tag (same GGUF, decision and calibration) plus the author's vision projector from the same pinned revision. PNGs given as images on /api/decide or --image on the command line, up to 16 per request, are read with the state and the questions. /v1/systemone keeps TypeSafe's text-only schema; the text tags keep their path. The runner loads libmtmd from the pinned llama.cpp release next to libllama, prepares the images and the state once per request and evaluates each question from that split, as stock llama-server tokenizes a multimodal prompt. libmtmd's prompt-bearing debug messages stay out of the log. The projector stays on the runner's device, with the CPU fallback of #47. Parity against stock llama-server b11146 with the projector on the same device: 21 requests and 65 questions, every decision the same, option logits within 9.6e-6 on an RTX 4070, 9.5e-6 on an RTX 5090 (the shipped 8,192-token context) and 1.2e-5 on an x86-64 CPU; the 505 text questions pass on both GPUs. docs/families/winnow.md has the numbers and the commands.
cobanov
pushed a commit
that referenced
this pull request
Oct 6, 2026
The Windows base install and desktop app take llama.cpp's win-vulkan-x64 build of the pinned b11146 instead of win-cpu: the same CPU backends plus ggml-vulkan.dll, so GGUF models use a GPU of any vendor on Windows. OLLAYA_DEVICE accepts vulkan and vulkan:<n>. Under auto, CUDA comes first when its pack is installed, then the discrete GPU with the most free memory; an integrated Vulkan GPU shares system memory and is not shown to beat the CPU, so it takes an explicit vulkan:<n> (CUDA's integrated devices, such as a DGX Spark's GB10, and Metal stay eligible). A GPU runner that fails while it loads starts again on the CPU (#47). Parity is per device (ADR 0003, point 7): against stock llama-server b11146 on the same Vulkan device, an Intel Arc 140T gives winnow:e4b 505/505 within 1.1e-5 and jevk5 593/593 within 9.5e-6, and an RTX 4090 gives winnow:e4b 505/505, jeb:9b 494/494 and cygnet:12b 502/502, all within 1.1e-5. On a Windows build of this branch with an RTX 4090 and a UHD 770, auto picks the RTX 4090 through Vulkan, the CPU when only the UHD 770 is visible, and the UHD 770 with vulkan:0. docs/measurements/intel-arc-140t-parity.md records the Arc runs and their distance to CUDA, which is about the CPU's.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A GGUF runner on
autoalready falls back to the CPU itself when llama.cpp reports an error while loading or warming up on the GPU. It can't do that when the GPU backend takes the process down before the runner answers, for example a driver fault or an abort inside ggml (as with the CUDA kernels in #42). The server then reported a failed load, and every request for that model failed the same way.The scheduler now handles GGUF models on
autothe same way it handles ONNX models with a GPU pack. If the runner doesn't come up, the scheduler starts it again with--device cpu, using the same executable and environment, and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), because then any GPU driver on the machine is in the path.Tested on Linux:
ollaya-servertests pass.