Skip to content

feat(windows): enable Vulkan for GGUF models - #27

Merged
cobanov merged 14 commits into
ollaya-dev:mainfrom
MauricioPerera:intel-arc-vulkan-pr
Oct 6, 2026
Merged

cobanov merged 14 commits into
ollaya-dev:mainfrom
MauricioPerera:intel-arc-vulkan-pr

Conversation

@MauricioPerera

@MauricioPerera MauricioPerera commented Sep 27, 2026 •

Copy link
Copy Markdown

CUDA gate: fails

The requested CUDA-to-Arc parity gate is not met. Against the reviewer's stock b11146 RTX 4090 fixture, stock Windows Vulkan and the Ollaya runner agree on 502/505 decisions, with maximum normalized option-logit difference 0.3261 (limit 1e-3). This PR is not ready to merge under that requirement.

Same-backend checks pass for Winnow 505/505 and JevK5 593/593; these establish that the runner reproduces stock Vulkan, not CUDA acceptance. Private backend experiments identify activation-quantization differences and improve agreement with a host mathematical reference, but every tested candidate still fails known CUDA counterexamples. No private arithmetic patch or modified backend binary is distributed by this PR.

An isolated same-input reproduction and maintainer request seek the actual CUDA result needed for further numerical localization. That diagnostic does not replace the 505-question fixture.

Changes

  • Bundle the pinned llama.cpp Windows Vulkan release in the base CLI and desktop packages, including ggml-vulkan.dll and the CPU backends.
  • Accept OLLAYA_DEVICE=vulkan and vulkan:<n> for GGUF runners, report the selected Vulkan device, and reject Vulkan requests for ONNX models.
  • Prefer CUDA, then discrete Vulkan, then integrated Vulkan for auto; retain CPU fallback when GPU loading or warm-up fails.
  • Document device selection, package contents, and the measured CUDA gate failure.
  • Correct export/replay fixture names to goldens-vulkan for Vulkan references.
  • Merge current main v0.7.5 (32acb6d) without conflicts.

Validation after the v0.7.5 merge

  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets --locked -- -D warnings
  • cargo test --workspace --locked
  • cargo build --release --locked -p ollaya
  • Release parity_llama example rebuilt.
  • npm run typecheck in site/
  • Stock Vulkan recheck with the rebuilt runner on four selected CUDA regression requests: 12/15 decisions, maximum option-logit difference 0.3261. This is not a new full 505-question run.

The Windows Vulkan archive SHA-256 is pinned to 55a378aa095b466979d85075234f66d7655c7a7483222af0c006c0e55b4d7bd6. Hardware measurements use Intel Core Ultra 9 285H / Intel Arc 140T, driver 32.0.101.8860. Other Vulkan drivers and hybrid CUDA/Vulkan hardware remain untested here.

Evidence

Full report, model pins, hashes, commands and numerical diagnostics.

@cobanov cobanov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you, this is well done: the device parsing, the CUDA-first/discrete-first auto_gpu with its test, and the docs. I built it (a release dry run of this PR merged on main) and measured it on Windows 11 with an RTX 4090 and an Intel UHD 770, and I can't merge it as it is, because of what Ollaya promises about answers.

Every model and device Ollaya uses has to pass the parity gate: the same decision on every test question, and option logits within 1e-3 of the reference (for GGUF models, stock llama-server of the same build on CUDA). parity_llama with winnow:e4b (505 questions):

backend on the RTX 4090 decisions option logits max probabilities max 5 questions, p50
CUDA (ggml-cuda.dll) 505/505 1.3e-5 3.0e-6 104 ms
Vulkan (ggml-vulkan.dll, this PR) 501/505 0.32 0.055 138 ms

GGML_VK_DISABLE_F16=1 gives the same Vulkan numbers. Because auto would pick Vulkan whenever the CUDA pack is missing, any Windows machine with a GPU would get answers that differ from every other platform, with no sign of it.

What changed on main in the meantime (#23): the Windows CUDA pack now ships llama.cpp's ggml-cuda.dll, so NVIDIA GPUs run GGUF models on Windows with the exact CUDA numbers above. That covers the NVIDIA part of this PR.

For Intel and AMD GPUs, Vulkan is still the right idea. It can go in once it passes the gate: for example with a llama.cpp build or Vulkan settings whose results match, or with evidence that the difference comes from the reference, not the backend. Your Arc numbers with parity_llama would be very useful (cargo run --release -p ollaya-runner --example parity_llama -- <model-dir> <goldens.jsonl> Vulkan0; goldens and the model dir layout are in docs/families/winnow.md). I'm keeping the PR open for that; thanks again for the work.

@cobanov

cobanov commented Sep 28, 2026

Copy link
Copy Markdown
Member

For completeness: the Intel UHD 770 in the same machine (Vulkan1, an integrated GPU like the Arc 140T) is outside the gate too, with 503 of the 508 checks failing, option logits off by up to 0.18.

@MauricioPerera

Copy link
Copy Markdown
Author

Thanks for the detailed review. I am preparing the same parity check on the Intel Arc 140T, with Winnow-E4B Q8_0 pinned to the catalog revision and llama.cpp b11146.

Could you share the exact CUDA reference goldens.jsonl and its metadata used for the 505-question gate (or a download URL)? I found the generation/replay tools and the published decision/calibration configs, but not the golden fixtures in the repository or the ollaya-dev/winnow Hub files. Using your fixture will let me report an exact comparison against the reference you tested, rather than substituting a Vulkan-generated reference.

I will also compare the runner against stock llama-server on the Arc to distinguish runtime/parity-plan issues from differences between the CUDA and Vulkan backends. I will update this PR with measured results; the earlier latency benchmarks alone do not establish the required parity.

@cobanov

cobanov commented Sep 28, 2026

Copy link
Copy Markdown
Member

Here is the fixture: https://gist.github.com/cobanov/c808f61976a8210235d8faf8ae91aa15

It is winnow-e4b-q8_0-goldens-cuda.jsonl (123 requests, 505 questions, 15 of them rejected). It was made by ollaya_convert.families.llm_common.export_llama, which sends the winnow-v1 reference prompt to a stock llama-server from llama.cpp b11146 with the CUDA backend, on an RTX 4090. The model is EldanRing/Winnow-E4B@734302fe gguf/Winnow-E4B-Q8_0.gguf. The gist also has the registry's decision.json and calibration.json, and the README gives the sha256 and the exact parity_llama command.

This file is the gate for Vulkan: compare the runner on the Arc against it. A stock llama-server on Vulkan also helps, because it separates a Vulkan-backend difference from a runner problem. It can't replace the CUDA reference, though, since both sides of that comparison would share the same backend.

@MauricioPerera

Copy link
Copy Markdown
Author

Thanks for the fixture. I verified its SHA-256 (b0b15b5646e8cda17364b6097a9613eb231a88a4813b52c129423313447695fe), the model SHA-256, and that the supplied decision/calibration files match the registry files used here.

The full gate fails on the Intel Arc 140T, driver 32.0.101.8860, stock Windows Vulkan b11146:

Comparison against your CUDA fixture Decisions Max option-logit difference Max probability difference
Ollaya Vulkan runner, all 505 questions 502/505 0.3261 0.05600
Stock Vulkan llama-server, all 505 questions 502/505 0.32614444 0.05600171

The 123 cases and 15 rejected requests match, including prompt IDs, split points, candidates, wire order and state metadata. The changed decisions are router/mask_text/domain, many_options_40/intent, and customer_service_000080/churn_risk.

For the separate same-backend check, Ollaya reproduces stock Vulkan: Winnow 505/505 (max logit 1.143e-5) and JevK5 593/593 (9.521e-6). These do not replace the CUDA gate; they show the observed difference is already present with the stock backend.

I tested the three changed-decision cases plus the largest-logit-error case (4 requests, 15 questions). Disabling F16, forcing MMVQ, disabling the optimization set, and disabling flash attention all fail. Disabling cooperative matrices reduces the maximum error to 0.1517 but still changes 2 decisions; combining that with forced MMVQ gives 0.1414 and changes 3 decisions. CPU matches all 15 selected decisions but also fails the logit tolerance (0.1915). No tested configuration passes.

The pinned source has precision differences worth investigating further: cooperative matrices force F32 activations through F16, and Vulkan's Q8_1 activation scales are stored in FP16 while CUDA MMQ uses FP32 scales for Q8_0. This is a source-level observation, not a validated numerical fix.

I pushed the main merge, corrected the export/replay tools so Vulkan fixtures are named goldens-vulkan instead of goldens-cpu, and recorded the commands, pins, hashes and measurements in the report. The distribution documentation now states that the measured CUDA gate fails. I have not raised the tolerance or treated native Vulkan parity as acceptance. There is no validated arithmetic correction in this PR, so it is not ready to ship Vulkan under this gate.

@MauricioPerera

Copy link
Copy Markdown
Author

Follow-up: I tested the Q8_0 MMQ activation-scale hypothesis in a private b11146 build. The experiment preserves FP32 scales through storage, shared memory and registers, and gives the quantization producer a separate pipeline identity for buffer reuse. MMVQ and formats requiring activation sums keep their original representation.

A local control built with the same MSVC 19.51 / shaderc 2026.3 tools reproduces the previous selected-case results. With cooperative matrices disabled, the four known CUDA regression requests give:

Build Decisions, selected 15 questions Max option-logit difference
Control 13/15 0.1517
FP32 scales 14/15 0.1337
Control, forced MMVQ 12/15 0.1414
FP32 scales, forced MMVQ 14/15 0.1428

Both experimental variants still change td/customer_service_000080/churn_risk. These counterexamples already fail the unchanged 1e-3 gate, so I did not expand the experiment to all 505 questions. The full stock result remains 502/505 with max error 0.3261.

The existing Q8_0/Q4_1 operation tests pass 24/24 supported cases in both builds, using the upstream operation-test tolerance. On the separate 50-question JevK5 CPU diagnostic fixture, the experiment changes agreement from 47/50 to 48/50 and max logit error from 0.2072 to 0.1716. That CPU comparison cannot establish CUDA parity.

Changing these activation scales alone is insufficient. The remaining error is not isolated to a specific operator, and this does not prove that a future backend correction is impossible. I pushed the measurements and result data with fixture/binary hashes in 43fed54. The private arithmetic patch and binaries are not part of this PR or the distribution. There is still no validated numerical correction, and the PR does not meet the CUDA gate.

@MauricioPerera

Copy link
Copy Markdown
Author

@cobanov I narrowed the local CPU/Arc difference to activation quantization and prepared a same-input standalone CUDA diagnostic.

With identical captured weights and inputs, the first Q8_0 multiplication differs by up to 0.234026 between local CPU and Arc. A GPU-buffer capture identifies 192 one-unit activation-quantization differences near halfway values; accumulating the GPU's actual quantized bytes against a Float64 mathematical reference leaves only 0.000107 maximum error. A private FMA-residual division refinement removes all 192 differences and makes the scales match the host on this input, but still fails your selected CUDA questions (13/15, max logit 0.1731; disabling fusion gives 14/15 and 0.1202). This is not a validated final-logit fix.

The pinned CUDA configuration uses -use_fast_math, so matching a host formula cannot establish the RTX 4090's actual result. Could you run the packet's matmul-replay on your stock b11146 CUDA backend and share comparison.json, the CUDA library/build identity, and preferably cuda-output.bin in a zip?

The packet fixes 64 input rows and includes hashed CPU/control-Arc/refined-Arc outputs. It extracts and hash-checks the Q8_0 weight tensor from the pinned author GGUF rather than embedding model weights. No CUDA source modifications are needed. I tested preparation, standalone compilation and CPU/Vulkan replay on Windows; the CPU replay exactly reproduces the corresponding captured rows. I also downloaded and hash-verified the published data. Linux/CUDA execution remains untested here.

This is an isolated raw-tensor diagnostic, not a replacement for the 505-question fixture or its 1e-3 gate. The full stock result remains a failure. The report and quantization data contain the measurements. I also merged current main v0.7.5 without conflicts; llama and configuration tests pass, and the release parity example rebuilds. The private backend experiments are not distributed by this PR.

@solarpush

Copy link
Copy Markdown

Hi @cobanov and @MauricioPerera,

First of all, thank you both for the great work on Ollaya and the impressive speed of execution on the Jev/Winnow-style models!

Following a similar motivation to Mauricio’s Vulkan PR, I have been working on adding AMD ROCm/HIP support to Ollaya (using the official llama.cpp libggml-hip.so / ggml-hip.dll release binaries from b11146). While Vulkan is a great generic path, ROCm on Linux and Windows is mature, officially packaged by ggml-org, and provides near-CUDA native performance on AMD GPUs (16.7s for 505 questions on a Radeon RX 9070, ~30 q/s).

I encountered the exact same parity dilemma on the Winnow-E4B fixture, and I wanted to share empirical measurements that shed light on Mauricio's findings regarding the CUDA gate.

1. The ~0.32–0.35 logit drift is universal across all non-CUDA backends

We evaluated the exact same 123 requests (505 questions, 15 rejected) from @cobanov's RTX 4090 reference (winnow-e4b-q8_0-goldens-cuda.jsonl) on an AMD Radeon RX 9070 (RDNA 4 / Navi 48 / gfx1201) and an AMD Ryzen 9 9950X (Linux CPU):

Platform & Backend Hardware Decisions vs CUDA Reference Max Option Logit Drift Max Proba Drift Time (505 q)
CUDA (Reference) NVIDIA RTX 4090 505 / 505 (100%) — — ~10 s
ROCm / HIP AMD Radeon RX 9070 500 / 505 (99.01%) 0.3298 0.0318 16.7 s
Vulkan (PR #27) Intel Arc 140T 502 / 505 (99.41%) 0.3261 0.0560 —
Vulkan (PR #27) NVIDIA RTX 4090 501 / 505 (99.21%) 0.3200 0.0550 14 s
CPU Linux (x86-64) AMD Ryzen 9 9950X 505 / 505 (100%) 0.3535 0.0443 228 s

Notice that even stock Linux x86-64 CPU exhibits a max logit difference of 0.3535 against the CUDA reference (failing 501/505 questions on the 1e-3 gate, despite 100% decision agreement).

This confirms Mauricio's localization: the drift does not stem from a defect in Vulkan or ROCm, but from NVIDIA's -use_fast_math SFU transcendental approximations compounding across the 32 transformer layers of Gemma 4 attention soft-capping ($30 \times \tanh(x/30)$). <= src Gemini.

2. Native Parity passes the Gate unconditionally (ADR 0003)

In Architecture Decision 0003, it is stated:

"The goldens come from the family's Python reference prompt, evaluated on the same pinned build's llama-server with the same plan, one goldens file per device."

When evaluated against a stock llama-server b11146 started on the device under test via llm_common.replay, Ollaya's in-process runner reproduces the reference with sub-$10^{-5}$ bit-level agreement:

  • AMD ROCm (ROCm0 on RX 9070):
    123 cases (15 rejected), 505 questions on rocm:0 in 16.7s:
    decisions 505/505 (100.00%), option logits max 1.246e-5 p99 1.132e-5, probabilities max 3.022e-6
    PASS
    
  • AMD Ryzen 9 9950X (cpu):
    123 cases (15 rejected), 505 questions on cpu in 233.6s:
    decisions 505/505 (100.00%), option logits max 1.168e-5 p99 1.093e-5, probabilities max 2.773e-6
    PASS
    

For comparison, CUDA on RTX 4090 was measured at 1.29e-5. The native runner-to-server fidelity on ROCm and CPU is identical to CUDA.

3. Fixtures Gist

I generated and published the turnkey fixtures for ROCm and CPU (including metadata and verification instructions) here:
👉 https://gist.github.com/solarpush/76f669b9b5121f1b26b13e550aeef378

4. Thoughts on the way forward

Cross-hardware parity on deep soft-capped models like Winnow-E4B cannot reach < 1e-3 across different GPU architectures unless fast-math is disabled during CUDA compilation (which is out of reach since Ollaya relies on stock ggml-org prebuilts).

As stated in ADR 0003, using per-device class golden references (goldens-cuda, goldens-rocm, goldens-cpu, and device-specific Vulkan) validates what Ollaya actually cares about: guaranteeing that the Rust in-process runner behaves identically to upstream llama-server on that machine. Alternatively, if a single cross-backend golden is desired, the logit tolerance for Winnow-E4B would need to be relaxed to ~0.35 — and even then, such an arbitrary threshold would remain a fragile moving target across diverse architectures
(e.g. linux-arm64 NEON/SVE, Apple Silicon Metal, or next-gen GPU ISAs) where non-linear accumulation across 32 layers can drift slightly differently without representing any actual runner regression.

I will open a clean, dedicated PR for the ROCm backend with all packaging, Windows/Linux loader support, and full measurements. Happy to coordinate with both of you!

# Conflicts:
#	crates/ollaya-runner/src/llama/mod.rs
#	scripts/llama-cpp.sh
#	site/docs/faq.md
@cobanov

cobanov commented Oct 1, 2026

Copy link
Copy Markdown
Member

@MauricioPerera, I owe you a correction and an apology: I reviewed this PR against a stricter rule than Ollaya applies to its own devices.

ADR 0003 sets the parity reference for a GGUF model as the same pinned llama.cpp build's llama-server on the same device: one goldens file per device. That is how the CPU path ships. Asking Vulkan on the Arc to match the RTX 4090's CUDA numbers within 1e-3 was a gate no other backend passes. I measured it today. On winnow:e4b, our own x86-64 CPU goldens agree with CUDA on 501 of 505 decisions, with option log-probabilities up to 0.28 apart; Vulkan on the same RTX 4090 gives 501/505 and 0.32. On cygnet:12b (Gemma 4 12B) the CPU gives 496/502 and 2.76, Vulkan 497/502 and 1.40. The difference you localized so carefully is how the backends round, not a problem in this PR. Sorry for the time this cost you.

By the ADR's rule this PR passes: your Arc 140T results (Winnow 505/505, JevK5 593/593 against stock Vulkan llama-server b11146), and mine on an RTX 4090 with this branch merged with main:

model decisions option logits max 5 questions p50, Vulkan CUDA
winnow:e4b 505/505 1.1e-5 149 ms 103 ms
jeb:9b 494/494 7.6e-6 250 ms 128 ms
cygnet:12b 502/502 7.7e-6 862 ms 213 ms

We'd like to merge it, with two changes:

  1. auto and integrated GPUs. Keep CUDA first and discrete Vulkan GPUs next, but leave integrated GPUs to an explicit OLLAYA_DEVICE=vulkan:<n> for now. An integrated GPU shares system memory, and nothing shows yet that it beats the CPU. If you can time parity_llama <model-dir> <goldens> Vulkan0 --latency against cpu --latency on the Arc 140T, and the Arc is clearly faster, we can let auto pick it in a follow-up.
  2. docs/measurements/. Please keep intel-arc-140t-parity.md, shortened to the setup, the same-backend parity table and the CUDA comparison as information. Drop the three JSON files from the private-build experiments, since they describe builds Ollaya doesn't ship. The FAQ sentence about the smoke test can go too; I'll update the docs with the per-device numbers.

On our side:

Thanks for sticking with this.

cobanov added a commit that referenced this pull request Oct 2, 2026
…n the GPU (#47)

A GGUF runner on `auto` already falls back to the CPU itself when llama.cpp reports an error while
loading or warming up on the GPU. It cannot when the GPU backend takes the process down before
the runner answers: a driver fault, or an abort inside ggml, as with the CUDA kernels of #42. The
server then reported a failed load, and every request for the model failed the same way.

The scheduler now treats GGUF models on `auto` as it treats ONNX models with a GPU pack: if the
runner does not come up, it starts it again with `--device cpu` (same executable and environment)
and logs why. This matters more once a GPU backend ships in the base install (Vulkan, #27), where
any GPU driver on the machine is in the path.

Tested: an HTTP test whose fake runner exits before answering unless it runs on the CPU loads on
the CPU (and fails without this change); a model that fails on every device still reports its
own error.

Co-authored-by: cobanov <29142615+cobanov@users.noreply.github.com>
cobanov added a commit that referenced this pull request Oct 2, 2026
…winnow, results/README

ADR 0003 gets a point 7: a device passes when the runtime matches stock llama-server on that same
device, and differences between devices are measured and published, not gated. It quotes the
cross-device numbers from the RTX 4090 run: Vulkan is no further from CUDA than the CPU backend
(winnow:e4b 0.32 against 0.28 in log-probability), and on its own goldens it passes (505 of 505).
This is the paragraph promised in #27.

docs/families/winnow.md no longer says Vulkan fails the gate: that compared it with CUDA's goldens.

results/README.md documents the data behind ollaya.dev/results: machines.json, the three run
suites and their fields, and how to add a run.
…eeps its parity tables

The two changes from the review of ollaya-dev#27:

- auto picks CUDA first, then the discrete GPU with the most free memory. An integrated Vulkan GPU
  shares system memory and nothing shows yet that it beats the CPU, so it now takes an explicit
  OLLAYA_DEVICE=vulkan:<n>. The filter is on Vulkan's integrated devices only: llama.cpp b11146
  reports CUDA's integrated devices (the GB10 of a DGX Spark) as IGPU too, and those keep CUDA, as
  does Metal, which reports a GPU. The test covers the three cases.
- docs/measurements/intel-arc-140t-parity.md keeps the setup, the same-backend parity table (the
  gate), the CUDA comparison as information and the reproduction; the three JSON files from the
  private-build experiments go, since they describe builds Ollaya doesn't ship.

docs/distribution.md gives Vulkan's per-device parity (Arc 140T, and an RTX 4090 through Vulkan
with its speed next to CUDA) under ADR 0003's rule, and the README, FAQ and API docs say that
auto uses a discrete Vulkan GPU on Windows. Also merged main (0.11.0) into the branch: the
conflicts were in scripts/llama-cpp.sh's header, docs/api.md (OLLAYA_THREADS) and the FAQ.

@cobanov cobanov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both review points are in (auto leaves integrated Vulkan GPUs to vulkan:, and the Arc 140T doc keeps its parity tables), main is merged, and a Windows build of this branch picks the RTX 4090 through Vulkan under auto, the CPU when only the UHD 770 is visible, and the UHD 770 with OLLAYA_DEVICE=vulkan:0.

@cobanov
cobanov merged commit 6dff202 into ollaya-dev:main Oct 6, 2026
17 checks passed
@cobanov

cobanov commented Oct 6, 2026

Copy link
Copy Markdown
Member

Merged, thanks @MauricioPerera, and thanks for the patience through the review. We made the two changes on your branch: auto leaves integrated Vulkan GPUs to an explicit vulkan:<n> (only Vulkan's: llama.cpp also reports a DGX Spark's CUDA GB10 as integrated, and that keeps CUDA), and the Arc 140T doc keeps the setup, the same-device parity table, the CUDA comparison as information and the reproduction. A Windows build of the branch picked the RTX 4090 through Vulkan under auto, the CPU when only the UHD 770 was visible, and the UHD 770 with OLLAYA_DEVICE=vulkan:0. It ships in the next release.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants