Skip to content

intel_xpu attention backend produces garbage output on Arc Pro B70 (BMG-G31) — persists after #196 fix and with its workaround #243

Description

@rahulunair

Summary

The intel_xpu attention backend produces garbage/incoherent output on Intel Arc Pro B70 (Battlemage BMG-G31) from the very first request. This persists on current sglang main including the #196 fix (PR sgl-project/sglang#23280, cu_seqlens_k_new=None at all call sites — verified present in our build), and also with the #196 workaround (--chunked-prefill-size -1 --disable-radix-cache), where the prepopulated-KV prefill path should not be hit at all. The same checkpoint on the same GPU is coherent with --attention-backend triton, and coherent under vLLM-XPU — so the model, driver, and runtime are fine; the intel_xpu kernel path is not.

Since #196 was validated on Arc B580 (BMG-G21), but seeing this on BMG-G31 i.e. a different bug than #196 with the same symptom.

Environment

  • GPU: Intel Arc Pro B70 32 GB, BMG-G31 (intel_gpu_bmg_g31), single card (ZE_AFFINITY_MASK=0)
  • Kernel: Linux 7.0.11 (CachyOS)
  • UMD: kobuk-team PPA — libze-intel-gpu1 26.18.38308.1-1~24.04~ppa1, libze1 1.28.2-1~24.04~ppa1 (same PPA the sgl-kernel-xpu CI Dockerfile uses)
  • torch: 2.11.0+xpu
  • sglang: main @ bb33594c1 (0.5.6.post3.dev6213+gbb33594c1) — includes PR #23280 (grep cu_seqlens_k_new .../xpu_backend.py shows =None at all 8 call sites)
  • sgl-kernel (xpu): built from sgl-kernel-xpu main, reports sgl_kernel 0.1.8
  • Model: Qwen/Qwen3-0.6B (also reproduced with Qwen3-8B), bf16, local safetensors

Reproduction

python -m sglang.launch_server \
  --model-path /models/Qwen3-0.6B --device xpu --attention-backend intel_xpu \
  --chunked-prefill-size -1 --disable-radix-cache \
  --host 0.0.0.0 --port 8200
curl -s http://127.0.0.1:8200/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"default","messages":[{"role":"user","content":"What is 2+2? Answer briefly."}],"max_tokens":256,"temperature":0}'

Output (first request after a fresh boot — no radix cache, no chunked prefill):

timesunci&actionieectvasive근受 time[]( TMPro动ineTransformίoltage tínhеждуectorётeliness[](我们要ANCED.tokenquivoroid群众inDMETHOD AppModuledfunding单ting内容uablyhipakah<void...

Server logs are clean (no errors; normal decode throughput ~56 tok/s).

Control experiments (same box, same checkpoint)

configuration output
--attention-backend intel_xpu (defaults) garbage
--attention-backend intel_xpu --chunked-prefill-size -1 --disable-radix-cache (the #196 workaround) garbage
--attention-backend triton (same sglang build) coherent (<think>\n\n</think>\n\n4)
vLLM XPU backend, same checkpoint, same GPU coherent
sglang 0.5.6.post3 release build (pre-#23280) + intel_xpu garbage (so not a recent regression)

Happy to run any diagnostic build/env permutation on this hardware

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions