Benchmarks
All rows come from one desktop. The quality column is a different test on different rows.
Swipe the table sideways for more columns →
| Date | Model | Quant | Hardware | Serving stack | Context | 1-stream tok/s | 8-conc tok/s | Quality |
|---|---|---|---|---|---|---|---|---|
| 2026-10-01 | Qwen3.8-27B | FP8 (official pre-quantized, block-wise e4m3, 128x128 scales) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce, host-staged for this eval) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 53 | 227 | correctness gate: 10/10 |
| 2026-10-01 | Qwen3.8-27B | AutoRound Int4 (community build, relabelled as GPTQ for the fast kernel path) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce, host-staged for this eval) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 73 | 317 | LiveCodeBench-9 (greedy) / exact-answer(12) / gate: 6/9 (vs FP8 7/9) / 10/12 / 10/10 |
| 2026-10-01 | Qwen3.8-27B | FP8 (online, at load, official BF16 weights) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | - | - | engine start time / KV pool (vision tower on vs. off): 226-239s / 396,601 tokens -> 131s / 452,706 tokens (+14%) |
| 2026-09-30 | Qwen3.8-27B | FP8 (online, at load, official BF16 weights) | 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce, P2P tripped) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 37 | 155 | - |
| 2026-09-30 | Qwen3.8-27B | FP8 (online, at load, official BF16 weights) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 52.9 | 251 | LiveCodeBench-24 (medium thinking) / websites (/75) / readiness (10 tasks): 14/24 / 75 / 9 pass, 0 fail, 1 n/a |
| 2026-09-29 | Qwen3.8-27B | GPTQ-Int4 (sym, g128) + int4 lm_head + int4 MTP layer | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 86.3 | 342 | HumanEval / GSM8K(250) / gate / exact-recall(160): 156/164 / 240 (n.s. vs 243) / 10/10 / 158 |
| 2026-09-29 | Qwen3.8-27B | GPTQ-Int4 (sym, g128, BF16 MTP head) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 67 | 311 | MTP speculative tokens (k): k=4 |
| 2026-09-28 | Qwen3.8-27B | GPTQ-Int4 (sym, g128, BF16 MTP head) | 1x Intel Arc Pro B60 24GB | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
81920 | 23.9 | 160.6 | correctness gate: 10/10 |
| 2026-09-28 | Qwen3.8-27B | GPTQ-Int4 (sym, g128, BF16 MTP head) | 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 13.7 | 102 | - |
| 2026-09-28 | Qwen3.8-27B | GPTQ-Int4 (sym, g128, BF16 MTP head) | 2x Intel Arc Pro B60 24GB (TP=2, host-staged allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 42.3 | 172.5 | correctness gate / tool call: 10/10 / pass |
| 2026-09-28 | Qwen3.8-27B | GPTQ-Int4 (sym, g128, BF16 MTP head) | 2x Intel Arc Pro B60 24GB (TP=2, P2P allreduce) | vLLM (intel/llm-scaler-vllm:0.26.0-b2) 0.26.1.dev0+g568afb3a1 |
262144 | 56 | 301.8 | correctness gate / tool call / stress: 10/10 / pass / 312 ok, 0 fail (900s) |
Context is the configured --max-model-len. Speeds are
greedy decode tokens/second, 1 stream and 8 concurrent. Hover a row
for notes.