Repository navigation
feat(windows): enable Vulkan for GGUF models - #27
Conversation
cobanov
left a comment
There was a problem hiding this comment.
Thank you, this is well done: the device parsing, the CUDA-first/discrete-first auto_gpu with its test, and the docs. I built it (a release dry run of this PR merged on main) and measured it on Windows 11 with an RTX 4090 and an Intel UHD 770, and I can't merge it as it is, because of what Ollaya promises about answers.
Every model and device Ollaya uses has to pass the parity gate: the same decision on every test question, and option logits within 1e-3 of the reference (for GGUF models, stock llama-server of the same build on CUDA). parity_llama with winnow:e4b (505 questions):
| backend on the RTX 4090 | decisions | option logits max | probabilities max | 5 questions, p50 |
|---|---|---|---|---|
CUDA (ggml-cuda.dll) |
505/505 | 1.3e-5 | 3.0e-6 | 104 ms |
Vulkan (ggml-vulkan.dll, this PR) |
501/505 | 0.32 | 0.055 | 138 ms |
GGML_VK_DISABLE_F16=1 gives the same Vulkan numbers. Because auto would pick Vulkan whenever the CUDA pack is missing, any Windows machine with a GPU would get answers that differ from every other platform, with no sign of it.
What changed on main in the meantime (#23): the Windows CUDA pack now ships llama.cpp's ggml-cuda.dll, so NVIDIA GPUs run GGUF models on Windows with the exact CUDA numbers above. That covers the NVIDIA part of this PR.
For Intel and AMD GPUs, Vulkan is still the right idea. It can go in once it passes the gate: for example with a llama.cpp build or Vulkan settings whose results match, or with evidence that the difference comes from the reference, not the backend. Your Arc numbers with parity_llama would be very useful (cargo run --release -p ollaya-runner --example parity_llama -- <model-dir> <goldens.jsonl> Vulkan0; goldens and the model dir layout are in docs/families/winnow.md). I'm keeping the PR open for that; thanks again for the work.
|
For completeness: the Intel UHD 770 in the same machine ( |
|
Thanks for the detailed review. I am preparing the same parity check on the Intel Arc 140T, with Winnow-E4B Q8_0 pinned to the catalog revision and llama.cpp b11146. Could you share the exact CUDA reference I will also compare the runner against stock llama-server on the Arc to distinguish runtime/parity-plan issues from differences between the CUDA and Vulkan backends. I will update this PR with measured results; the earlier latency benchmarks alone do not establish the required parity. |
|
Here is the fixture: https://gist.github.com/cobanov/c808f61976a8210235d8faf8ae91aa15 It is This file is the gate for Vulkan: compare the runner on the Arc against it. A stock llama-server on Vulkan also helps, because it separates a Vulkan-backend difference from a runner problem. It can't replace the CUDA reference, though, since both sides of that comparison would share the same backend. |
|
Thanks for the fixture. I verified its SHA-256 ( The full gate fails on the Intel Arc 140T, driver 32.0.101.8860, stock Windows Vulkan b11146:
The 123 cases and 15 rejected requests match, including prompt IDs, split points, candidates, wire order and state metadata. The changed decisions are router/mask_text/domain, many_options_40/intent, and customer_service_000080/churn_risk. For the separate same-backend check, Ollaya reproduces stock Vulkan: Winnow 505/505 (max logit 1.143e-5) and JevK5 593/593 (9.521e-6). These do not replace the CUDA gate; they show the observed difference is already present with the stock backend. I tested the three changed-decision cases plus the largest-logit-error case (4 requests, 15 questions). Disabling F16, forcing MMVQ, disabling the optimization set, and disabling flash attention all fail. Disabling cooperative matrices reduces the maximum error to 0.1517 but still changes 2 decisions; combining that with forced MMVQ gives 0.1414 and changes 3 decisions. CPU matches all 15 selected decisions but also fails the logit tolerance (0.1915). No tested configuration passes. The pinned source has precision differences worth investigating further: cooperative matrices force F32 activations through F16, and Vulkan's Q8_1 activation scales are stored in FP16 while CUDA MMQ uses FP32 scales for Q8_0. This is a source-level observation, not a validated numerical fix. I pushed the main merge, corrected the export/replay tools so Vulkan fixtures are named |
|
Follow-up: I tested the Q8_0 MMQ activation-scale hypothesis in a private b11146 build. The experiment preserves FP32 scales through storage, shared memory and registers, and gives the quantization producer a separate pipeline identity for buffer reuse. MMVQ and formats requiring activation sums keep their original representation. A local control built with the same MSVC 19.51 / shaderc 2026.3 tools reproduces the previous selected-case results. With cooperative matrices disabled, the four known CUDA regression requests give:
Both experimental variants still change The existing Q8_0/Q4_1 operation tests pass 24/24 supported cases in both builds, using the upstream operation-test tolerance. On the separate 50-question JevK5 CPU diagnostic fixture, the experiment changes agreement from 47/50 to 48/50 and max logit error from 0.2072 to 0.1716. That CPU comparison cannot establish CUDA parity. Changing these activation scales alone is insufficient. The remaining error is not isolated to a specific operator, and this does not prove that a future backend correction is impossible. I pushed the measurements and result data with fixture/binary hashes in |
|
@cobanov I narrowed the local CPU/Arc difference to activation quantization and prepared a same-input standalone CUDA diagnostic. With identical captured weights and inputs, the first Q8_0 multiplication differs by up to 0.234026 between local CPU and Arc. A GPU-buffer capture identifies 192 one-unit activation-quantization differences near halfway values; accumulating the GPU's actual quantized bytes against a Float64 mathematical reference leaves only 0.000107 maximum error. A private FMA-residual division refinement removes all 192 differences and makes the scales match the host on this input, but still fails your selected CUDA questions (13/15, max logit 0.1731; disabling fusion gives 14/15 and 0.1202). This is not a validated final-logit fix. The pinned CUDA configuration uses The packet fixes 64 input rows and includes hashed CPU/control-Arc/refined-Arc outputs. It extracts and hash-checks the Q8_0 weight tensor from the pinned author GGUF rather than embedding model weights. No CUDA source modifications are needed. I tested preparation, standalone compilation and CPU/Vulkan replay on Windows; the CPU replay exactly reproduces the corresponding captured rows. I also downloaded and hash-verified the published data. Linux/CUDA execution remains untested here. This is an isolated raw-tensor diagnostic, not a replacement for the 505-question fixture or its 1e-3 gate. The full stock result remains a failure. The report and quantization data contain the measurements. I also merged current main v0.7.5 without conflicts; llama and configuration tests pass, and the release parity example rebuilds. The private backend experiments are not distributed by this PR. |
|
Hi @cobanov and @MauricioPerera, First of all, thank you both for the great work on Ollaya and the impressive speed of execution on the Jev/Winnow-style models! Following a similar motivation to Mauricio’s Vulkan PR, I have been working on adding AMD ROCm/HIP support to Ollaya (using the official llama.cpp I encountered the exact same parity dilemma on the Winnow-E4B fixture, and I wanted to share empirical measurements that shed light on Mauricio's findings regarding the CUDA gate. 1. The ~0.32–0.35 logit drift is universal across all non-CUDA backendsWe evaluated the exact same 123 requests (505 questions, 15 rejected) from
Notice that even stock Linux x86-64 CPU exhibits a max logit difference of 0.3535 against the CUDA reference (failing 501/505 questions on the 1e-3 gate, despite 100% decision agreement). This confirms Mauricio's localization: the drift does not stem from a defect in Vulkan or ROCm, but from NVIDIA's 2. Native Parity passes the Gate unconditionally (ADR 0003)In Architecture Decision 0003, it is stated:
When evaluated against a stock
For comparison, CUDA on RTX 4090 was measured at 3. Fixtures GistI generated and published the turnkey fixtures for ROCm and CPU (including metadata and verification instructions) here: 4. Thoughts on the way forwardCross-hardware parity on deep soft-capped models like Winnow-E4B cannot reach As stated in ADR 0003, using per-device class golden references ( I will open a clean, dedicated PR for the ROCm backend with all packaging, Windows/Linux loader support, and full measurements. Happy to coordinate with both of you! |
# Conflicts: # crates/ollaya-runner/src/llama/mod.rs # scripts/llama-cpp.sh # site/docs/faq.md
|
@MauricioPerera, I owe you a correction and an apology: I reviewed this PR against a stricter rule than Ollaya applies to its own devices. ADR 0003 sets the parity reference for a GGUF model as the same pinned llama.cpp build's By the ADR's rule this PR passes: your Arc 140T results (Winnow 505/505, JevK5 593/593 against stock Vulkan
We'd like to merge it, with two changes:
On our side:
Thanks for sticking with this. |
…n the GPU (#47) A GGUF runner on `auto` already falls back to the CPU itself when llama.cpp reports an error while loading or warming up on the GPU. It cannot when the GPU backend takes the process down before the runner answers: a driver fault, or an abort inside ggml, as with the CUDA kernels of #42. The server then reported a failed load, and every request for the model failed the same way. The scheduler now treats GGUF models on `auto` as it treats ONNX models with a GPU pack: if the runner does not come up, it starts it again with `--device cpu` (same executable and environment) and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), where any GPU driver on the machine is in the path. Tested: an HTTP test whose fake runner exits before answering unless it runs on the CPU loads on the CPU (and fails without this change); a model that fails on every device still reports its own error. Co-authored-by: cobanov <29142615+cobanov@users.noreply.github.com>
…winnow, results/README ADR 0003 gets a point 7: a device passes when the runtime matches stock llama-server on that same device, and differences between devices are measured and published, not gated. It quotes the cross-device numbers from the RTX 4090 run: Vulkan is no further from CUDA than the CPU backend (winnow:e4b 0.32 against 0.28 in log-probability), and on its own goldens it passes (505 of 505). This is the paragraph promised in #27. docs/families/winnow.md no longer says Vulkan fails the gate: that compared it with CUDA's goldens. results/README.md documents the data behind ollaya.dev/results: machines.json, the three run suites and their fields, and how to add a run.
…eeps its parity tables The two changes from the review of ollaya-dev#27: - auto picks CUDA first, then the discrete GPU with the most free memory. An integrated Vulkan GPU shares system memory and nothing shows yet that it beats the CPU, so it now takes an explicit OLLAYA_DEVICE=vulkan:<n>. The filter is on Vulkan's integrated devices only: llama.cpp b11146 reports CUDA's integrated devices (the GB10 of a DGX Spark) as IGPU too, and those keep CUDA, as does Metal, which reports a GPU. The test covers the three cases. - docs/measurements/intel-arc-140t-parity.md keeps the setup, the same-backend parity table (the gate), the CUDA comparison as information and the reproduction; the three JSON files from the private-build experiments go, since they describe builds Ollaya doesn't ship. docs/distribution.md gives Vulkan's per-device parity (Arc 140T, and an RTX 4090 through Vulkan with its speed next to CUDA) under ADR 0003's rule, and the README, FAQ and API docs say that auto uses a discrete Vulkan GPU on Windows. Also merged main (0.11.0) into the branch: the conflicts were in scripts/llama-cpp.sh's header, docs/api.md (OLLAYA_THREADS) and the FAQ.
cobanov
left a comment
There was a problem hiding this comment.
Both review points are in (auto leaves integrated Vulkan GPUs to vulkan:, and the Arc 140T doc keeps its parity tables), main is merged, and a Windows build of this branch picks the RTX 4090 through Vulkan under auto, the CPU when only the UHD 770 is visible, and the UHD 770 with OLLAYA_DEVICE=vulkan:0.
|
Merged, thanks @MauricioPerera, and thanks for the patience through the review. We made the two changes on your branch: |
CUDA gate: fails
The requested CUDA-to-Arc parity gate is not met. Against the reviewer's stock b11146 RTX 4090 fixture, stock Windows Vulkan and the Ollaya runner agree on 502/505 decisions, with maximum normalized option-logit difference 0.3261 (limit 1e-3). This PR is not ready to merge under that requirement.
Same-backend checks pass for Winnow 505/505 and JevK5 593/593; these establish that the runner reproduces stock Vulkan, not CUDA acceptance. Private backend experiments identify activation-quantization differences and improve agreement with a host mathematical reference, but every tested candidate still fails known CUDA counterexamples. No private arithmetic patch or modified backend binary is distributed by this PR.
An isolated same-input reproduction and maintainer request seek the actual CUDA result needed for further numerical localization. That diagnostic does not replace the 505-question fixture.
Changes
ggml-vulkan.dlland the CPU backends.OLLAYA_DEVICE=vulkanandvulkan:<n>for GGUF runners, report the selected Vulkan device, and reject Vulkan requests for ONNX models.auto; retain CPU fallback when GPU loading or warm-up fails.goldens-vulkanfor Vulkan references.32acb6d) without conflicts.Validation after the v0.7.5 merge
cargo fmt --all --checkcargo clippy --workspace --all-targets --locked -- -D warningscargo test --workspace --lockedcargo build --release --locked -p ollayaparity_llamaexample rebuilt.npm run typecheckinsite/The Windows Vulkan archive SHA-256 is pinned to
55a378aa095b466979d85075234f66d7655c7a7483222af0c006c0e55b4d7bd6. Hardware measurements use Intel Core Ultra 9 285H / Intel Arc 140T, driver 32.0.101.8860. Other Vulkan drivers and hybrid CUDA/Vulkan hardware remain untested here.Evidence
Full report, model pins, hashes, commands and numerical diagnostics.