Skip to content

feat: add AMD ROCm/HIP support on Linux and Windows - #33

Open
solarpush wants to merge 1 commit into
ollaya-dev:mainfrom
solarpush:feat/rocm-support
Open

solarpush wants to merge 1 commit into
ollaya-dev:mainfrom
solarpush:feat/rocm-support

Conversation

@solarpush

Copy link
Copy Markdown
- Runner & Server: - Add rocm and rocm-dynamic features with ONNX Runtime ROCm EP. - Add llama.cpp HIP backend support with case-insensitive device matching. - Implement seamless dynamic fallback to CPU runner on GPU/HIP failure or OOM. - Support OLLAYA_DEVICE=rocm and isolated pack loading in lib/ollaya/rocm/.

- Installation & Detection:
  - Detect AMD GPUs (vendor 0x1002, /dev/kfd, /dev/dri) in install.sh and install.ps1.
  - Decode sysfs KFD gfx_target_version (e.g. gfx1201, gfx1100, gfx942) without rocminfo. - Guarantee transparent CPU fallback if driver or GPU is unsupported.

- Packaging & CI: - Add --rocm packaging flag and bundle llama.cpp b11146 HIP backend. - Add multi-stage Docker build (rocm/dev-ubuntu-24.04:6.3) and CI/CD matrix jobs.

@solarpush

Copy link
Copy Markdown
Author

Hi @cobanov,

Here is a comprehensive overview of this PR, detailing the architecture choices, the automated hardware detection, and the parity verification results on AMD RDNA hardware.


1. Motivation & Architecture

AMD GPUs have strong native support in llama.cpp via the official HIP backend (libggml-hip.so on Linux, ggml-hip.dll on Windows), distributed directly as release assets by ggml-org for build b11146.

This PR integrates ROCm/HIP into Ollaya following the exact design principles established for CUDA and Metal:

  1. Isolated Pack Loading (lib/ollaya/rocm/):

    • The ROCm libraries and runner (target/rocm-dynamic/release/ollaya) are packaged into an optional --rocm pack (ollaya-<platform>-rocm.tar.zst).
    • The base CPU build and daemon carry zero hard dependencies on ROCm/HIP runtimes.
    • When present, the daemon dynamically discovers and loads the ROCm pack next to the install prefix.
  2. Automated Hardware Detection (Zero External Dependencies):

    • scripts/install.sh: Detects AMD GPUs by probing /dev/kfd and reading sysfs (/sys/class/kfd/kfd/topology/nodes/*/properties/gfx_target_version). It matches known AMD targets (gfx1201 for RDNA 4, gfx1100/gfx1101/gfx1102 for RDNA 3, gfx1030 for RDNA 2, gfx942 for CDNA 3, etc.) without requiring rocminfo or external utilities.
    • scripts/install.ps1: Detects AMD display adapters on Windows.
    • Transparent fallback: If an AMD card is present but the driver/kernel is too old, the installer gracefully defaults to the CPU build.
  3. Robust Runtime Device Selection & Fallback:

    • Supports OLLAYA_DEVICE=rocm and rocm:<n>.
    • auto_gpu ranks devices rationally: CUDA > ROCm > discrete GPU > integrated GPU > CPU.
    • If GPU VRAM is exhausted or HIP fails during model warm-up, the daemon falls back dynamically and cleanly to the CPU runner without crashing or aborting requests.

2. Parity Gate Validation (ADR 0003 Compliance)

In Architecture Decision 0003, the parity reference rule is specified as:

"The goldens come from the family's Python reference prompt, evaluated on the same pinned build's llama-server with the same plan, one goldens file per device."

We evaluated the complete test set on Winnow-E4B Q8_0 (123 requests, 15 rejected, 505 questions) on an AMD Radeon RX 9070 (RDNA 4 Navi 48 / gfx1201) and an AMD Ryzen 9 9950X (16 cores, AVX-512):

A. Native ROCm Parity (parity_llama vs stock llama-server b11146):

123 cases (15 rejected), 505 questions on rocm:0 in 16.7s:
decisions 505/505 (100.00%), option logits max 1.246e-5 p99 1.132e-5, probabilities max 3.022e-6 p99 2.530e-6
PASS

B. Native CPU Parity (parity_llama vs stock llama-server b11146):

123 cases (15 rejected), 505 questions on cpu in 233.6s:
decisions 505/505 (100.00%), option logits max 1.168e-5 p99 1.093e-5, probabilities max 2.773e-6 p99 2.496e-6
PASS
Backend Hardware Decisions Max Option Logit Drift Max Proba Drift Time (505 q) Gate Status
CUDA NVIDIA RTX 4090 505 / 505 (100%) $1.290 \times 10^{-5}$ $3.000 \times 10^{-6}$ ~10 s PASS
ROCm AMD Radeon RX 9070 505 / 505 (100%) $1.246 \times 10^{-5}$ $3.022 \times 10^{-6}$ 16.7 s (~30.2 q/s) PASS
CPU AMD Ryzen 9 9950X 505 / 505 (100%) $1.168 \times 10^{-5}$ $2.773 \times 10^{-6}$ 233.6 s PASS

The native runner-to-server fidelity on ROCm matches CUDA to sub-$10^{-5}$ precision, well within the strict $10^{-3}$ threshold. ROCm provides a 13.6x speedup over a 16-core CPU.


3. Empirical Insight on Cross-Backend Drift & CUDA Fast-Math

For completeness, we also compared the raw outputs of all non-CUDA backends directly against the RTX 4090 golden reference (winnow-e4b-q8_0-goldens-cuda.jsonl):

Backend & Hardware Decisions vs CUDA Reference Max Option Logit Drift Max Proba Drift
ROCm (AMD Radeon RX 9070) 500 / 505 (99.01%) 0.3298 0.0318
Vulkan (Intel Arc 140T, PR #27) 502 / 505 (99.41%) 0.3261 0.0560
Vulkan (NVIDIA RTX 4090, PR #27) 501 / 505 (99.21%) 0.3200 0.0550
CPU Linux (AMD Ryzen 9 9950X) 505 / 505 (100.00%) 0.3535 0.0443

Key observations:

  1. Even standard Linux x86-64 CPU exhibits a max logit difference of 0.3535 against the CUDA reference (failing 501/505 questions on the 1e-3 gate).
  2. The ~0.32–0.35 drift observed across Vulkan, ROCm, and CPU is universal: it stems from NVIDIA's -use_fast_math SFU approximations compounding across the 32 transformer layers of Gemma 4 attention soft-capping ($30 \times \tanh(x/30)$).
  3. Attempting to force a single cross-backend golden by relaxing the tolerance to ~0.35 would remain a fragile moving target across diverse architectures (e.g. linux-arm64 NEON/SVE, Apple Silicon Metal, or next-gen GPU ISAs) where non-linear accumulation across 32 layers can drift slightly differently without representing any actual runner regression.
  4. This empirically confirms why ADR 0003's per-device native reference model ("one goldens file per device") is the sound, robust way to verify runner integrity.

4. Turnkey Golden Fixtures Gist

The complete verified fixtures, JSON metadata, and reproduction instructions have been published in this Gist:
👉 https://gist.github.com/solarpush/76f669b9b5121f1b26b13e550aeef378

To verify on an AMD machine:

cargo run --release -p ollaya-runner --example parity_llama -- \
  <model-dir> winnow-e4b-q8_0-goldens-rocm.jsonl auto

5. Verification Checklist

All workspace checks pass cleanly:

  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets --locked -- -D warnings
  • cargo clippy -p ollaya --features ollaya-runner/rocm-dynamic --all-targets --locked -- -D warnings
  • cargo test --workspace --locked (144 / 144 unit and integration tests pass)
  • cargo build --release --locked -p ollaya
  • In-process parity verified: 505/505 PASS ($1.25 \times 10^{-5}$) on Radeon RX 9070

    - Runner & Server:
      - Add rocm and rocm-dynamic features with ONNX Runtime ROCm EP.
      - Add llama.cpp HIP backend support with case-insensitive device matching.
      - Implement seamless dynamic fallback to CPU runner on GPU/HIP failure or OOM.
      - Support OLLAYA_DEVICE=rocm and isolated pack loading in lib/ollaya/rocm/.

    - Installation & Detection:
      - Detect AMD GPUs (vendor 0x1002, /dev/kfd, /dev/dri) in install.sh and install.ps1.
      - Decode sysfs KFD gfx_target_version (e.g. gfx1201, gfx1100, gfx942) without rocminfo.
      - Guarantee transparent CPU fallback if driver or GPU is unsupported.

    - Packaging & CI:
      - Add --rocm packaging flag and bundle llama.cpp b11146 HIP backend.
      - Add multi-stage Docker build (rocm/dev-ubuntu-24.04:6.3) and CI/CD matrix jobs.
@solarpush

solarpush commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

Hello, could you let me know if you agree to proceed with this approach?

  • device-based parity
  • addition of ROCm (and oneAPI for Intel Arc in the future, if necessary)

If so, I will update the PR to include the necessary libraries for ORT (currently missing; even though the function call is correct, the library integration in package.sh is missing only gguf works on rocm gpu now ).

edit :
Alternatively, I actually recommend dropping ROCm support for ONNX entirely for now.
│ MIGraphX/ORT ROCm is simply not mature enough for dynamic LLM graphs: kernel fusion
│ and rocBLAS optimizations fail on dynamic shapes, leading to graph fragmentation and
│ constant CPU/GPU fallbacks with terrible performance.
│ We can keep ROCm focused exclusively on GGUF / llama.cpp (HIP), where it is stable
and
│ genuinely fast, and let ONNX models run on CPU on AMD hardware.

i wait your merge on vulkan to rebase change and use auto_gpu fn on #27

@cobanov

cobanov commented Oct 2, 2026

Copy link
Copy Markdown
Member

@solarpush, sorry for the slow reply, and thanks for the detailed write-up.

Yes to both:

  • Parity per device. That is the rule for every device. ADR 0003 now says it explicitly (point 7): a GPU passes when the runtime on it matches stock llama-server of the pinned build on that same GPU. feat(windows): enable Vulkan for GGUF models #27 (Vulkan) is going in on that basis.
  • AMD through llama.cpp's HIP backend. Welcome.

Two scope changes before we can merge:

  1. GGUF only. Please drop the ONNX Runtime ROCm part. ONNX Runtime removed the ROCm execution provider in 1.23 and points to MIGraphX instead (ROCm EP docs), so that path has no future here. On AMD machines the ONNX families keep running on the CPU.
  2. No Docker image yet. Please leave the ROCm Docker build and its CI jobs out of this PR. We can add an image once the runtime part has shipped and people have used it.

For the merge itself:

  • Goldens. Generate them from stock llama-server (b11146, HIP build) on your GPU, and post the parity_llama output against them: at least winnow:e4b, and jeb:9b if it fits. The gate is identical prompt ids, every decision the same, and option logits within 1e-3.
  • Rebase. Please rebase on main; the PR conflicts now.

@solarpush

Copy link
Copy Markdown
Author

@cobanov
Thanks,
I solve it this month and push fix with complete rebase.

@cobanov

cobanov commented Oct 6, 2026

Copy link
Copy Markdown
Member

@solarpush, a status update so your rebase starts from the right place: main has moved, and two things change for this PR.

  • Vulkan is merged (feat(windows): enable Vulkan for GGUF models #27). On Windows, GGUF models now use any vendor's discrete GPU through llama.cpp's Vulkan build in the base install, so AMD cards on Windows already have a GPU path. ROCm/HIP is still welcome as the faster path, and on Linux, where we ship no Vulkan build, it is the only one.
  • Device selection lives in crates/ollaya-runner/src/llama/mod.rs (auto_gpu): CUDA first, then the discrete GPU with the most free memory; OLLAYA_DEVICE takes cuda:<n> and vulkan:<n>. A HIP device should slot in there (rocm:<n> or hip:<n>), with its own pack directory next to cuda_v13.

What we need to merge, unchanged from before:

  1. GGUF only (no ONNX Runtime ROCm), and no Docker image or CI jobs in this PR.
  2. Parity on your RX 9070: stock llama-server of the pinned build (b11146, ggml-org's HIP release) on that GPU as the reference, and parity_llama against it for at least winnow:e4b (505 questions): every decision the same, option logits within 1e-3. docs/families/winnow.md has the commands; the Vulkan doc in docs/measurements/intel-arc-140t-parity.md shows the shape we keep.

We have no AMD GPU to run it ourselves, so the numbers have to come from your machine. Thanks!

@solarpush

Copy link
Copy Markdown
Author

Yes,
Thanks, i'm start to works on next week
I have change not commited yet after inspection of migraphX for onnx.
Busy before next monday.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants