Benchmark sweep harness for local model inference on SM120.
My real traffic (non-bench) across v4-flash and v4.1-flash for codex, claude code and agent evals.
My practical use of v4-flash started with 0731 release as original(preview) 0424 wasn't as usable day-to-day for my main session (non-subagent) coding workflows. 0424 had common occurances of reasoning (at max) degenerating after 300-400k context mark, a fraction of invalid tool call could be recovered with mechanical fixes, but reasoning repetition and mandarin on long context effectively constrained the usable context below 300k.
Sglang deployment for dsv4-flash-0731 also took a series of fixes to decrease error rates on tool calls and stabilize the caching for OAI's /v1/responses and Anthropics /v1/messages which broke caching at times as claude code updates started to use new features available in /v1/messages (like mid-coversation system messages), ref fix/dsv4-longctx-production-overlay, while the sglang deployment was changed from my custom W4A4 stack to upstream SGLang of
b03ac355
with W4A8 MoE that beat the previous custom W4A4 stack in all matched benchmarks.
Upgrade to dsv4.1-flash saw substantail performance gains in prefill and decode but needed caching fixes, 256GB RAM L2 HiCache enabled to exceeed previous dsv4-flash-0731 96% avarage. Mean per-request decode went 48.8 → 97.9 tok/s. Building on 0xSero's SM120 recipe, we added SWA-boundary retention and tuned 256 GB RAM HiCache to 67% SWA / 33% FULL KV by bytes. That RAM cache contributed additional 3.62%, bringing combined L1 + L2 coverage to 97.82%. The Level1Techs post provides the original comparison's; the V4.1 numbers below are updated to latest launch (L16):
| Metric | DSV4-FLASH-0731 | DSV4.1-FLASH |
|---|---|---|
| Duration | ||
| Reporting period (UTC) | Aug 24 – Sep 15, 2026 | Sep 20–25, 2026 (L16) |
| Wall-clock observation span | 531.70 h | 111.11 h |
| Sum of request-activity windows | 511.73 h | 111.03 h |
| Completed non-health generations | 256,325 | 70,477 |
| Health-check generations | 6,303 | 4,700 |
| POST /v1/messages | 175,668 | 62,615 |
| POST /v1/chat/completions | 62,436 | 7,100 |
| POST /v1/responses | 18,662 | 869 |
| POST /v1/messages/count_tokens | 2,356 | 685 |
| Prompt size, tokens | ||
| Average full prompt | 111,499 | 123,933 |
| Median prompt | 102,047 | 118,872 |
| p95 prompt | 264,293 | 249,896 |
| Largest prompt | 473,107 | 272,426 |
| Average cached prompt tokens | 107,688 | 121,229 |
| Average uncached prompt tokens | 3,811 | 2,704 |
| Latency, seconds — p50 / p95 / p99 | ||
| Queue / admission | 0.64 / 2.08 / 27.14 | 0.92 / 3.06 / 16.75 |
| TTFT, engine-side | 1.21 / 7.09 / 42.00 | 1.31 / 6.31 / 23.29 |
| Prefill | 0.27 / 2.27 / 18.09 | 0.24 / 1.67 / 8.51 |
| Decode | 5.04 / 49.89 / 152.32 | 3.64 / 31.49 / 83.17 |
| E2E, engine-side | 6.47 / 58.60 / 161.35 | 5.34 / 36.28 / 87.64 |
| Throughput, average | ||
| Per-request decode, tok/s | 48.8 | 97.9 |
| Batch generation, tok/s | 137.4 | 216.9 |
| Prefill input, tok/s | 2,444 | 2,552 |
| DSPARK(5) acceptance length / rate | N/A — not enabled | 3.74 / 0.548 |
| Tokens | ||
| Input tokens | 28,579,867,197 | 8,734,391,441 |
| Output tokens | 148,540,907 | 53,386,522 |
| Of which reported reasoning | 87,042,866 | 33,351,924 |
| Average output per generation | 580 | 758 |
| Cache & health | ||
| Token-weighted cache hit | 96.58% | 97.82% |
| Input served from RAM L2 | 0% — disabled | 3.62% · 315,928,576 tokens |
| Zero-hit generations | 11,578 (4.52%) | 3,003 (4.26%) |
| HTTP 200 / 503 / 400, all routes | 303,043 / 40 / 37 | 117,009 / 4 / 23 |
| Other HTTP statuses | 404: 54, 500: 10 | 0 |
| HTTP 2xx rate, all routes | 99.953% | 99.977% |
| HTTP 2xx rate, inference routes only | 99.982% | 99.967% |
| Allocator OOM-warning log lines | 0 | 113 |
Generation, token and prompt rows use logged non-health completions; HTTP rows have a separate counting scope. See L16 sources, counter reconciliation and run outcome, deployment configuration and benchmark reports.
2026-08-20 · TP4 · 300 W/GPU. Pinned SGLang b03ac355 with native W4A8 MoE beat the previous custom W4A4 stack in all nine matched benchmark cells (8K/64K input, 1K output), improving output throughput and lowering TTFT, TPOT and end-to-end latency.
| Single-stream output throughput | Previous custom stack | SGLang main b03ac355 |
|---|---|---|
| 8K input / 1K output | 46.94 tok/s | 82.65 tok/s |
| 64K input / 1K output | 18.38 tok/s | 47.19 tok/s |
See the benchmark report
and Docker image (2026.08.0-cu130-sm120a).
- GPUs: 4x NVIDIA RTX 6000 Blackwell Pro Max-Q Workstation Edition (96GB x4)
- CPU: AMD Ryzen Threadripper PRO 7985WX (64-core)
- RAM: 512 GB DDR5 ECC (8x 64 GB Kingston KSM56R46BD4PMI-64HAI)
- Platform: ASUS Pro WS WRX90E-SAGE SE
- PSU: Super Flower Leadex Titanium 1700W ATX 3.1
- OS: Ubuntu 24.04 LTS
| Model | Released | Architecture | Params (total / active) | Context | Engine | Weights | Quantization | Config | Tput 2K·c64 (tok/s) | Tput 64K·c16 (tok/s) | SWE-bench Verified (mini-swe-agent v2.4.2) | 1M decode (tok/s, c1)¹ | 1M prefill TTFT (c1)¹ | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MiniMax-M3 | Jun 1, 2026 | MoE (GQA + sparse) | 428B / 23B | 1,048,576 | sglang | olka-fi/MiniMax-M3-MXFP4 | MXFP4 experts + MXFP8 linears | launch-bwrap-highconc.sh | 1,045 | 35 | 74.8% (374/500) | 39.2 | 7.7 s | native MXFP4 W4A4 experts (clamped SwiGLU-OAI) + MXFP8 weight-only linears + SM120 Triton MSA block-sparse attention; split-K MXFP8; TP4 |
| DeepSeek-V4-Flash | Apr 24, 2026 | MoE (MLA + sparse) | 284B / 13B | 1,048,576 | sglang | deepseek-ai/DeepSeek-V4-Flash | MXFP4 experts + FP8 rest (native) | sglang-single.yaml | 756 | 40 | 76.0% (380/500) | 11.1 | 6.5 s | native MXFP4 W4A4 experts + block-FP8 attention/dense + HMMA tensor-core sparse decode and prefill + split-KV long-context indexer (sm_120); TP4 |
| Qwen3.5-397B-A17B | Feb 16, 2026 | MoE | 397B / 17B | 262,144 | vLLM | nvidia/Qwen3.5-397B-A17B-NVFP4 | NVFP4 | vllm.yaml | 1,124 | 102 | — | — | — |
Models above are at 300 W, TP4, devstral-2 and minimax-m2.5 at diff w, tp are excluded, but you can find their results under bench
¹ Single-stream (c1) decode tok/s / prefill TTFT at 1047552-token input (= 1048576 − 1024).
bench: 128 random prompts per run, 1024 output tokens, input lengths from 2K to 64K.
One row per model × input length, at the highest concurrency run on both for that input. mem-BW = mean memory-bandwidth util; PCIe = mean aggregate PCIE tx/rx across 4 GPUs; GPU0–3 = mean per-GPU temp @ power draw; sys W = system power mean/peak.
| Model | Input | Conc | tok/s | mem-BW | PCIe tx/rx | GPU0 | GPU1 | GPU2 | GPU3 | sys W (mean/peak) |
|---|---|---|---|---|---|---|---|---|---|---|
| DSV4 | 2,048 | c80 | 697 | 30% | 33/41 GB/s | 89°/240W | 82°/251W | 87°/249W | 85°/262W | 1,003 / 1,099 |
| DSV4 | 4,096 | c48 | 435 | 29% | 32/41 GB/s | 88°/238W | 83°/250W | 87°/247W | 85°/260W | 994 / 1,166 |
| DSV4 | 8,192 | c56 | 248 | 27% | 33/43 GB/s | 89°/234W | 84°/247W | 87°/243W | 85°/255W | 979 / 1,173 |
| DSV4 | 16,384 | c48 | 149 | 24% | 35/47 GB/s | 88°/225W | 82°/237W | 87°/234W | 85°/246W | 942 / 1,155 |
| DSV4 | 32,768 | c32 | 77 | 22% | 34/44 GB/s | 87°/219W | 79°/228W | 86°/228W | 82°/237W | 913 / 1,009 |
| DSV4 | 65,536 | c24 | 39 | 19% | 34/45 GB/s | 88°/221W | 81°/232W | 86°/231W | 84°/242W | 926 / 1,054 |
| M3 | 2,048 | c80 | 1,076 | 50% | 19/18 GB/s | 89°/284W | 85°/280W | 89°/275W | 86°/265W | 1,104 / 1,184 |
| M3 | 4,096 | c48 | 382 | 50% | 29/27 GB/s | 90°/291W | 85°/285W | 89°/279W | 86°/269W | 1,124 / 1,191 |
| M3 | 8,192 | c56 | 237 | 49% | 29/27 GB/s | 91°/289W | 86°/286W | 89°/279W | 86°/269W | 1,123 / 1,189 |
| M3 | 16,384 | c48 | 129 | 47% | 30/26 GB/s | 91°/286W | 86°/283W | 89°/277W | 86°/267W | 1,113 / 1,201 |
| M3 | 32,768 | c32 | 68 | 44% | 30/25 GB/s | 90°/290W | 85°/285W | 89°/277W | 86°/267W | 1,119 / 1,201 |
| M3 | 65,536 | c24 | 35 | 43% | 29/24 GB/s | 91°/282W | 86°/282W | 89°/275W | 86°/264W | 1,102 / 1,201 |
Throughput — M3 wins only at 2K (1,076 vs 697); from 4K up DSV4 leads. mem-BW — M3 runs
at ~2× DSV4's mem-bandwidth util at every length (8K: 49% vs 27%): its dense MXFP8 linears are BW-bound, which
draws the ~150–200 W more per run. PCIe — DSV4 moves more interconnect traffic (33/41 vs 19/18 GB/s at 2K;
peaks ~100/131), its MLA all-to-all vs M3's lighter block-sparse collective. Per-GPU thermals show a stable
~7 °C chassis gradient (slots 0/2 hottest, slot 1 coolest) due to 4 GPUs being stacked tight with GPU0/GPU2 being sandwitched in the middle. VRAM is constant at 344 GB (DSV4, mem-fraction 0.90) / 317 GB (M3,
0.95). SM/memory clock rates weren't sampled by the harness, so they're not shown.
![]() |
![]() |
Output throughput vs concurrency across input lengths — DeepSeek-V4-Flash (left) and MiniMax-M3 (right). M3's short-context curves keep rising past the shared ceiling (2K peak 1,593 @c128); DSV4 saturates KV earlier.
Single-stream, input 32K → 1,047,552 (= 1M − 1024, one prompt per length, output 1024) — both at full 1M
context on the same 300 W box, so the comparison is power-controlled and the difference is purely architectural
(M3 GQA block-sparse vs DSV4 MLA). DSV4 = its best long-ctx config (split-KV indexer,
SGLANG_SM120_INDEXER_SPLIT); M3 at mem-fraction-static 0.95 (fp8 KV pool holds 1,069,580 tokens; the 1M run
holds 370 GB VRAM, KV 98%):
| Input | M3 TTFT | M3 decode tok/s | DSV4 TTFT | DSV4 decode tok/s | M3/DSV4 decode |
|---|---|---|---|---|---|
| 32,768 | 0.4 s | 72.6 | 0.31 s | 52.9 | 1.4× |
| 131,072 | 0.9 s | 65.8 | 0.79 s | 38.7 | 1.7× |
| 262,144 | 1.9 s | 59.3 | 1.58 s | 28.6 | 2.1× |
| 524,288 | 4.2 s | 50.1 | 3.26 s | 18.7 | 2.7× |
| ~1.04M | 7.7 s | 39.2 | 6.46 s | 11.1 | 3.5× |
Decode = steady-state 1000/TPOT (both models). M3 decode degrades only −46 % over a 32× context increase (72.6 → 39.2 tok/s); DSV4 drops −79 % even with its split-KV indexer (52.9 → 11.1). The gap widens with length (1.4× at 32K → 3.5× at 1M): M3's block-sparse decode scans only the top-16 KV blocks, its indexer auto-shards across up to 256 CTAs at c1, fp8 KV keeps KV bandwidth low, and split-K keeps the MXFP8 linears off the critical path. DSV4's decode is sparse-attention top-k too (TPOT ~17 ms single-user), but its 32K/64K KV saturates at 100 %, so its concurrent peak comes earlier. M3 prefill sustains ~130 k tok/s across all lengths (see deploy doc) — the 1M prompt is processed in 7.7 s.
![]() |
![]() |
1,047,552-token long sequence run, 300 W cap per GPU. DSV4 900 W mean / 1,055 W peak; M3 1,083 W / 1,196 W — M3 sustains ~183 W more, while decoding ~39 vs ~10 tok/s.
uv pip install -e .Requires a running inference server (OpenAI-compatible API on http://127.0.0.1:8000) and vllm
installed in the harness venv — bench-sweep drives vllm bench serve --backend openai, which works
against any OpenAI-compatible endpoint (vLLM or sglang). It invokes the bench tool as
python -m vllm.entrypoints.cli.main bench serve (so a stale vllm console script doesn't block
runs); override with --bench-cmd if needed.
Both sglang models — DeepSeek-V4-Flash and MiniMax-M3 — are served by our SM120 image
ambientlight/sglang-sm120-mxfp4. Build + details: docker/sm120-unified/.
# 1. Serve (server metrics on by default)
docker run --rm --gpus all --ipc=host -p 8000:8000 -e MODEL=dsv4 \
-v /mnt/hot/ambientlight/models/DeepSeek-V4-Flash:/model:ro \
ambientlight/sglang-sm120-mxfp4:latest # wait for /v1/models (~8 min)
# MiniMax-M3: -e MODEL=m3 -v /mnt/hot/ambientlight/models/minimax-m3-mxfp4:/model:ro
# 2. Matrix sweep against the :8000 endpoint (output 1024, step-size 8 → c1,2,4,8,16,24,…,128)
bench-sweep --matrix --telemetry \
--model-id deepseek-v4-flash --watt 300 \
--tokenizer /mnt/hot/ambientlight/models/DeepSeek-V4-Flash \
--input-lens 2048,4096,8192,16384,32768,65536 --output-len 1024 \
--step-size 8 --num-prompts 128 --max-error-rate 0.1# Full matrix sweep with telemetry
bench-sweep --matrix --telemetry \
--model-id qwen35-397b-a17b-nvfp4 \
--tokenizer /path/to/tokenizer \
--watt 250 \
--input-lens 2048,4096,8192,16384,32768,65536 \
--output-len 1024 \
--step-size 8 \
--num-prompts 128
# Per-input-len max concurrency caps (avoid OOM on large inputs)
bench-sweep --matrix --telemetry \
--model-id qwen35-397b-a17b-nvfp4 \
--tokenizer /path/to/tokenizer \
--watt 250 \
--input-lens 2048,4096,8192,16384,32768,65536 \
--max-concurrency 128,96,96,96,48,48 \
--output-len 1024
# Re-plot existing results without re-running benchmarks
bench-sweep --matrix --plot-only \
--model-id qwen35-397b-a17b-nvfp4 --watt 250 \
--input-lens 2048,4096,8192,16384,32768,65536 --output-len 1024
# Single concurrency sweep
bench-sweep --model-id qwen35-397b-a17b-nvfp4 \
--tokenizer /path/to/tokenizer --watt 250
# Dry run (print commands only)
bench-sweep --dry-run --model-id qwen35-397b-a17b-nvfp4 \
--tokenizer /path/to/tokenizer --watt 250Raw results use the model/power/TP/engine layout below, with named subdirectories
for comparison variants. See the V4-0731 results
and V4.1 benchmark reports.
Report collections without a consistently recorded power limit omit W{watt}.
bench/
{model}_W{watt}_TP{tp}_{engine}/
vllm.yaml | sglang.yaml # Server configuration
b.log # bench-sweep invocation command
bench_sweep.log # Full execution log (all runs)
{model}_random_{in}in_{out}out_c{concurrency}_W{watt}/
openai-infqps-concurrency{N}-{model}-{YYYYMMDD-HHMMSS}.json
telemetry.csv
telemetry_summary.json
telemetry_power.png
telemetry_kv_cache.png
plots/
{model}_{in}in_{out}out_W{watt}/
overview.png # Combined dashboard
throughput_vs_concurrency.png
ttft_vs_concurrency.png
tpot_vs_concurrency.png
itl_vs_concurrency.png
e2el_vs_concurrency.png
duration_vs_concurrency.png
{model}_compare_W{watt}/
compare_overview_p50.png
compare_overview_p95_p99.png
compare_throughput_vs_concurrency.png
compare_ttft_p50_vs_concurrency.png
compare_tpot_p50_vs_concurrency.png
compare_itl_p50_vs_concurrency.png
compare_e2el_p50_vs_concurrency.png
compare_duration_vs_concurrency.png
compare_efficiency_vs_concurrency.png # tok/s per watt
compare_power_vs_concurrency.png
compare_gpu_util_vs_concurrency.png
compare_mem_bw_util_vs_concurrency.png
compare_kv_cache_vs_concurrency.png
One file per (model, input_len, output_len, concurrency) run. Produced by vLLM's benchmark_serving.py. See upstream for the full schema.
Time-series GPU metrics sampled at ~2.4 Hz during each benchmark run. 4-GPU system (gpu0 through gpu3), with per-GPU columns repeated.
Header (30+ columns):
| Column | Type | Unit | Description |
|---|---|---|---|
timestamp |
float | Unix epoch (s) | Absolute timestamp |
elapsed_s |
float | seconds | Time since benchmark start |
gpu{N}_power_w |
float | Watts | power draw |
gpu{N}_mem_used_gb |
float | GB | Memory used |
gpu{N}_util_pct |
int | % | Compute utilization (0-100) |
gpu{N}_mem_bw_util_pct |
int | % | Memory bandwidth utilization |
gpu{N}_temp_c |
int | C | Temperature |
gpu{N}_pcie_tx_mb_s |
float | MB/s | PCIe transmit throughput |
gpu{N}_pcie_rx_mb_s |
float | MB/s | PCIe receive throughput |
kv_cache_pct |
float | % | KV cache utilization (from server /metrics) |
requests_running |
int | count | Active inference requests |
requests_waiting |
int | count | Queued requests |
Where {N} is 0, 1, 2, 3. Columns repeat for each GPU.
Example row (abbreviated):
1774517244.043,0.000,128.0,81.59,0,0,84,0.5,0.4,111.4,81.63,...,0.0,0,0
Pre-aggregated statistics from telemetry.csv per run (power, GPU util, KV cache, PCIe, etc.). Each per-run directory contains one.
vLLM models:
- Qwen3.5-397B-A17B — checkpoint: nvidia/Qwen3.5-397B-A17B-NVFP4
- MiniMax-M2.5 — checkpoint: lukealonso/MiniMax-M2.5-NVFP4
- Devstral-2-123B — checkpoint: mistralai/Devstral-2-123B-Instruct-2512, manually quantized to NVFP4 using LLM Compressor with
transformersv5 (one-shot calibration on nvidia/OpenCodeInstruct, 128 samples, 8192 seq len) — quantization script
served via:
PYTORCH_ALLOC_CONF=expandable_segments:True vllm serve /path/to/model --config /path/to/model/vllm.yaml --port 8000 -O3sglang models — both served by ambientlight/sglang-sm120-mxfp4 (-e MODEL={dsv4|m3}; see docker/sm120-unified/):
- DeepSeek-V4-Flash — checkpoint: deepseek-ai/DeepSeek-V4-Flash (native MXFP4 W4A4); native MXFP4 fused MoE + HMMA sparse-attention on SM120. Serve with
-e MODEL=dsv4. - MiniMax-M3 — checkpoint: olka-fi/MiniMax-M3-MXFP4 (community MXFP4 experts + MXFP8 linears); MXFP4 fused MoE (clamped SwiGLU-OAI) + MXFP8 split-K linears + SM120 Triton MSA block-sparse attention. Serve with
-e MODEL=m3.
| File | Content |
|---|---|
b.log |
Single-line bench-sweep invocation command with all CLI args |
bench_sweep.log |
Full concatenated stdout from all benchmark runs (config dumps, progress, result tables) |
sweep-single.log / bench_sweep_single.*.log |
Single-sequence (c1) long-context sweep logs — the 32K→1M runs behind the "Long-context scaling to 1M" table (M3: sweep-single.log) |



