Benchmarks

All rows come from one desktop. The quality column is a different test on different rows.

Swipe the table sideways for more columns →

Date Model Quant Hardware Serving stack Context 1-stream tok/s 8-conc tok/s Quality
2026-10-01 Qwen3.8-27B FP8 (official pre-quantized, block-wise e4m3, 128x128 scales) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce, host-staged for this eval) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 53 227 correctness gate: 10/10
2026-10-01 Qwen3.8-27B AutoRound Int4 (community build, relabelled as GPTQ for the fast kernel path) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce, host-staged for this eval) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 73 317 LiveCodeBench-9 (greedy) / exact-answer(12) / gate: 6/9 (vs FP8 7/9) / 10/12 / 10/10
2026-10-01 Qwen3.8-27B FP8 (online, at load, official BF16 weights) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 - - engine start time / KV pool (vision tower on vs. off): 226-239s / 396,601 tokens -> 131s / 452,706 tokens (+14%)
2026-09-30 Qwen3.8-27B FP8 (online, at load, official BF16 weights) 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce, P2P tripped) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 37 155 -
2026-09-30 Qwen3.8-27B FP8 (online, at load, official BF16 weights) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 52.9 251 LiveCodeBench-24 (medium thinking) / websites (/75) / readiness (10 tasks): 14/24 / 75 / 9 pass, 0 fail, 1 n/a
2026-09-29 Qwen3.8-27B GPTQ-Int4 (sym, g128) + int4 lm_head + int4 MTP layer 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 86.3 342 HumanEval / GSM8K(250) / gate / exact-recall(160): 156/164 / 240 (n.s. vs 243) / 10/10 / 158
2026-09-29 Qwen3.8-27B GPTQ-Int4 (sym, g128, BF16 MTP head) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 67 311 MTP speculative tokens (k): k=4
2026-09-28 Qwen3.8-27B GPTQ-Int4 (sym, g128, BF16 MTP head) 1x Intel Arc Pro B60 24GB vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
81920 23.9 160.6 correctness gate: 10/10
2026-09-28 Qwen3.8-27B GPTQ-Int4 (sym, g128, BF16 MTP head) 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 13.7 102 -
2026-09-28 Qwen3.8-27B GPTQ-Int4 (sym, g128, BF16 MTP head) 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 42.3 172.5 correctness gate / tool call: 10/10 / pass
2026-09-28 Qwen3.8-27B GPTQ-Int4 (sym, g128, BF16 MTP head) 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) vLLM (intel/llm-scaler-vllm:0.26.0-b2)
0.26.1.dev0+g568afb3a1
262144 56 301.8 correctness gate / tool call / stress: 10/10 / pass / 312 ok, 0 fail (900s)

Context is the configured --max-model-len. Speeds are greedy decode tokens/second, 1 stream and 8 concurrent. Hover a row for notes.