Running a 27B model on two Arc Pro B60s

Two Intel Arc Pro B60s in a normal desktop now serve Qwen3.8-27B at about 53 tokens/s for a single user and around 250 tokens/s across eight, with a 262K context.

The setup

  • 2× Arc Pro B60, 24GB each, in a consumer Alder Lake board. No PCIe switch.
  • Qwen3.8-27B, split across both cards (tensor parallel, -tp 2).
  • Intel's llm-scaler-vllm image, 0.26.0-b2 (vLLM 0.26.1).

Two cards were slower than one

First result: 13.7 tok/s on two cards, against about 24 on one. The kernel won't allow GPU-to-GPU transfers on this chipset (it isn't on the p2pdma allow-list), so every sync between the cards went through system RAM. Worse, that path can't be captured in a graph, so the whole model ran in slow eager mode.

The fix was piecewise graph capture with the all-reduce as a split point, plus MTP speculative decoding with the drafter kept out of the graph. That got it to 42 tok/s single and 172 at eight concurrent.

Turning on P2P anyway

The allow-list says "unknown", not "broken", so I wrote a small kernel module that allows this one host bridge. The catch: every GPU buffer then got mirrored into system RAM, 15–17GB per card, and the box ran out of memory in seconds. That turned out to be Intel's Level Zero runtime making allocations resident on the other card. Two environment variables fixed it.

Through system RAMP2P
Time to first token, 206K prompt509 s151 s
Throughput, 8 users147 tok/s260 tok/s

A watchdog switches back to the slower path if a card ever faults. It has done that twice in production, cleanly both times.

The slow leak

After a few hours, speed would drop from about 21 to 4 tok/s and the engine would eventually run out of memory. Intel's oneCCL library keeps a staging buffer for every message size it sees and never frees them, and real prompts come in a lot of sizes. Turning the cache off halved the speed. Rounding every message up to a fixed set of sizes fixed it: a two-hour soak of real agent traffic grew memory by 0.08GB.

FP8 or int4?

int4 is faster. FP8 solved more problems in a 24-problem LiveCodeBench run done as real agent tasks, mostly on the hard ones:

1 user8 usersLiveCodeBench (24)
int486 tok/s342 tok/s12
FP853 tok/s251 tok/s14

Every problem int4 solved, FP8 solved too. I went with FP8. A later search through community int4 builds found one that tied FP8 on a smaller test, but not convincingly enough to switch. I'm building my own int4 next.

Faster starts

The model ships with a vision encoder that vLLM loads and warms up on every start. I send images to a separate, smaller vision model, so I turned it off here (--language-model-only). Starts went from about 230 seconds to 131, and the KV cache grew 14%.

If you're picking a second card, check whether the board allows P2P first. Without it two cards still beat one, but long prompts and multi-user throughput suffer badly.