Summary
The intel_xpu attention backend produces garbage/incoherent output on Intel Arc Pro B70 (Battlemage BMG-G31) from the very first request. This persists on current sglang main including the #196 fix (PR sgl-project/sglang#23280, cu_seqlens_k_new=None at all call sites — verified present in our build), and also with the #196 workaround (--chunked-prefill-size -1 --disable-radix-cache), where the prepopulated-KV prefill path should not be hit at all. The same checkpoint on the same GPU is coherent with --attention-backend triton, and coherent under vLLM-XPU — so the model, driver, and runtime are fine; the intel_xpu kernel path is not.
Since #196 was validated on Arc B580 (BMG-G21), but seeing this on BMG-G31 i.e. a different bug than #196 with the same symptom.
Environment
- GPU: Intel Arc Pro B70 32 GB, BMG-G31 (
intel_gpu_bmg_g31), single card (ZE_AFFINITY_MASK=0)
- Kernel: Linux 7.0.11 (CachyOS)
- UMD: kobuk-team PPA —
libze-intel-gpu1 26.18.38308.1-1~24.04~ppa1, libze1 1.28.2-1~24.04~ppa1 (same PPA the sgl-kernel-xpu CI Dockerfile uses)
- torch:
2.11.0+xpu
- sglang: main @
bb33594c1 (0.5.6.post3.dev6213+gbb33594c1) — includes PR #23280 (grep cu_seqlens_k_new .../xpu_backend.py shows =None at all 8 call sites)
- sgl-kernel (xpu): built from sgl-kernel-xpu main, reports
sgl_kernel 0.1.8
- Model: Qwen/Qwen3-0.6B (also reproduced with Qwen3-8B), bf16, local safetensors
Reproduction
python -m sglang.launch_server \
--model-path /models/Qwen3-0.6B --device xpu --attention-backend intel_xpu \
--chunked-prefill-size -1 --disable-radix-cache \
--host 0.0.0.0 --port 8200
curl -s http://127.0.0.1:8200/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"default","messages":[{"role":"user","content":"What is 2+2? Answer briefly."}],"max_tokens":256,"temperature":0}'
Output (first request after a fresh boot — no radix cache, no chunked prefill):
timesunci&actionieectvasive근受 time[]( TMPro动ineTransformίoltage tínhеждуectorётeliness[](我们要ANCED.tokenquivoroid群众inDMETHOD AppModuledfunding单ting内容uablyhipakah<void...
Server logs are clean (no errors; normal decode throughput ~56 tok/s).
Control experiments (same box, same checkpoint)
| configuration |
output |
--attention-backend intel_xpu (defaults) |
garbage |
--attention-backend intel_xpu --chunked-prefill-size -1 --disable-radix-cache (the #196 workaround) |
garbage |
--attention-backend triton (same sglang build) |
coherent (<think>\n\n</think>\n\n4) |
| vLLM XPU backend, same checkpoint, same GPU |
coherent |
| sglang 0.5.6.post3 release build (pre-#23280) + intel_xpu |
garbage (so not a recent regression) |
Happy to run any diagnostic build/env permutation on this hardware
Summary
The
intel_xpuattention backend produces garbage/incoherent output on Intel Arc Pro B70 (Battlemage BMG-G31) from the very first request. This persists on currentsglangmain including the #196 fix (PR sgl-project/sglang#23280,cu_seqlens_k_new=Noneat all call sites — verified present in our build), and also with the #196 workaround (--chunked-prefill-size -1 --disable-radix-cache), where the prepopulated-KV prefill path should not be hit at all. The same checkpoint on the same GPU is coherent with--attention-backend triton, and coherent under vLLM-XPU — so the model, driver, and runtime are fine; theintel_xpukernel path is not.Since #196 was validated on Arc B580 (BMG-G21), but seeing this on BMG-G31 i.e. a different bug than #196 with the same symptom.
Environment
intel_gpu_bmg_g31), single card (ZE_AFFINITY_MASK=0)libze-intel-gpu1 26.18.38308.1-1~24.04~ppa1,libze1 1.28.2-1~24.04~ppa1(same PPA the sgl-kernel-xpu CI Dockerfile uses)2.11.0+xpubb33594c1(0.5.6.post3.dev6213+gbb33594c1) — includes PR #23280 (grep cu_seqlens_k_new .../xpu_backend.pyshows=Noneat all 8 call sites)sgl_kernel 0.1.8Reproduction
Output (first request after a fresh boot — no radix cache, no chunked prefill):
Server logs are clean (no errors; normal decode throughput ~56 tok/s).
Control experiments (same box, same checkpoint)
--attention-backend intel_xpu(defaults)--attention-backend intel_xpu --chunked-prefill-size -1 --disable-radix-cache(the #196 workaround)--attention-backend triton(same sglang build)<think>\n\n</think>\n\n4)Happy to run any diagnostic build/env permutation on this hardware