Repository navigation
Conversation
solarpush
commented
Sep 28, 2026
|
Hi @cobanov, Here is a comprehensive overview of this PR, detailing the architecture choices, the automated hardware detection, and the parity verification results on AMD RDNA hardware. 1. Motivation & ArchitectureAMD GPUs have strong native support in llama.cpp via the official HIP backend ( This PR integrates ROCm/HIP into Ollaya following the exact design principles established for CUDA and Metal:
2. Parity Gate Validation (ADR 0003 Compliance)In Architecture Decision 0003, the parity reference rule is specified as:
We evaluated the complete test set on Winnow-E4B Q8_0 (123 requests, 15 rejected, 505 questions) on an AMD Radeon RX 9070 (RDNA 4 Navi 48 / A. Native ROCm Parity (
|
| Backend | Hardware | Decisions | Max Option Logit Drift | Max Proba Drift | Time (505 q) | Gate Status |
|---|---|---|---|---|---|---|
| CUDA | NVIDIA RTX 4090 | 505 / 505 (100%) | ~10 s | PASS | ||
| ROCm | AMD Radeon RX 9070 | 505 / 505 (100%) | 16.7 s (~30.2 q/s) | PASS | ||
| CPU | AMD Ryzen 9 9950X | 505 / 505 (100%) | 233.6 s | PASS |
The native runner-to-server fidelity on ROCm matches CUDA to sub-$10^{-5}$ precision, well within the strict
3. Empirical Insight on Cross-Backend Drift & CUDA Fast-Math
For completeness, we also compared the raw outputs of all non-CUDA backends directly against the RTX 4090 golden reference (winnow-e4b-q8_0-goldens-cuda.jsonl):
| Backend & Hardware | Decisions vs CUDA Reference | Max Option Logit Drift | Max Proba Drift |
|---|---|---|---|
| ROCm (AMD Radeon RX 9070) | 500 / 505 (99.01%) | 0.3298 | 0.0318 |
| Vulkan (Intel Arc 140T, PR #27) | 502 / 505 (99.41%) | 0.3261 | 0.0560 |
| Vulkan (NVIDIA RTX 4090, PR #27) | 501 / 505 (99.21%) | 0.3200 | 0.0550 |
| CPU Linux (AMD Ryzen 9 9950X) | 505 / 505 (100.00%) | 0.3535 | 0.0443 |
Key observations:
- Even standard Linux x86-64 CPU exhibits a max logit difference of 0.3535 against the CUDA reference (failing 501/505 questions on the 1e-3 gate).
- The ~0.32–0.35 drift observed across Vulkan, ROCm, and CPU is universal: it stems from NVIDIA's
-use_fast_mathSFU approximations compounding across the 32 transformer layers of Gemma 4 attention soft-capping ($30 \times \tanh(x/30)$). - Attempting to force a single cross-backend golden by relaxing the tolerance to ~0.35 would remain a fragile moving target across diverse architectures (e.g.
linux-arm64NEON/SVE, Apple Silicon Metal, or next-gen GPU ISAs) where non-linear accumulation across 32 layers can drift slightly differently without representing any actual runner regression. - This empirically confirms why ADR 0003's per-device native reference model ("one goldens file per device") is the sound, robust way to verify runner integrity.
4. Turnkey Golden Fixtures Gist
The complete verified fixtures, JSON metadata, and reproduction instructions have been published in this Gist:
👉 https://gist.github.com/solarpush/76f669b9b5121f1b26b13e550aeef378
To verify on an AMD machine:
cargo run --release -p ollaya-runner --example parity_llama -- \
<model-dir> winnow-e4b-q8_0-goldens-rocm.jsonl auto5. Verification Checklist
All workspace checks pass cleanly:
-
cargo fmt --all --check -
cargo clippy --workspace --all-targets --locked -- -D warnings -
cargo clippy -p ollaya --features ollaya-runner/rocm-dynamic --all-targets --locked -- -D warnings -
cargo test --workspace --locked(144 / 144 unit and integration tests pass) -
cargo build --release --locked -p ollaya - In-process parity verified: 505/505 PASS (
$1.25 \times 10^{-5}$ ) on Radeon RX 9070
- Runner & Server:
- Add rocm and rocm-dynamic features with ONNX Runtime ROCm EP.
- Add llama.cpp HIP backend support with case-insensitive device matching.
- Implement seamless dynamic fallback to CPU runner on GPU/HIP failure or OOM.
- Support OLLAYA_DEVICE=rocm and isolated pack loading in lib/ollaya/rocm/.
- Installation & Detection:
- Detect AMD GPUs (vendor 0x1002, /dev/kfd, /dev/dri) in install.sh and install.ps1.
- Decode sysfs KFD gfx_target_version (e.g. gfx1201, gfx1100, gfx942) without rocminfo.
- Guarantee transparent CPU fallback if driver or GPU is unsupported.
- Packaging & CI:
- Add --rocm packaging flag and bundle llama.cpp b11146 HIP backend.
- Add multi-stage Docker build (rocm/dev-ubuntu-24.04:6.3) and CI/CD matrix jobs.
bf56fee to
f312c74
Compare
|
Hello, could you let me know if you agree to proceed with this approach?
If so, I will update the PR to include the necessary libraries for ORT (currently missing; even though the function call is correct, the library integration in edit : i wait your merge on vulkan to rebase change and use auto_gpu fn on #27 |
|
@solarpush, sorry for the slow reply, and thanks for the detailed write-up. Yes to both:
Two scope changes before we can merge:
For the merge itself:
|
|
@cobanov |
|
@solarpush, a status update so your rebase starts from the right place: main has moved, and two things change for this PR.
What we need to merge, unchanged from before:
We have no AMD GPU to run it ourselves, so the numbers have to come from your machine. Thanks! |
|
Yes, |