Two Arc Pro B60s for vLLM: what hung and what fixed it

Intel's documented way to serve a model on two Arc Pro B60s works. For one user it gives me 39.2 tokens/s on Qwen3.8-27B (int4). Turn on XPU graphs and speculative decoding and that becomes 73.7, but with 32 GB of RAM about one start in four hung, with both cards resetting their copy engines.

The tl;dr: I went from 32 GB to 64 GB of RAM and 45 starts on kernel 6.17 came up without a hang. Everything else below is how I got there and what didn't matter.

The box

  • MSI X470 Gaming Plus Max, Ryzen 5 5600, 64 GB DDR4-3200 (it was 32 GB)
  • 2× Arc Pro B60 24 GB, each at PCIe 3.0 x8, peer-to-peer through the CPU
  • Ubuntu 24.04.4, Intel's offline installer 26.18.8.2: kernel 6.17.0-1009-intel, GuC 70.60.0
  • intel/llm-scaler-vllm:0.26.0-b2, tensor parallel 2

I run the kernel and GuC firmware from the installer, 6.17.0-1009-intel and 70.60.0. The newer ones I tried are further down.

The config

This is what runs now, with my paths swapped out. The two cache mounts matter more than they look (see below). The decode numbers after it are from the int4 checkpoint with the same graph, MTP and oneCCL settings.

docker run --name vllm --privileged --device /dev/dri:/dev/dri \
  -v /dev/dri/by-path:/dev/dri/by-path:ro \
  -v /srv/models:/models:ro \
  -v /srv/cache/vllm-compile:/root/.cache/vllm \
  -v /srv/cache/triton:/root/.triton \
  -p 8000:8000 --shm-size 32g --ipc=host \
  -e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
  -e VLLM_XPU_ENABLE_XPU_GRAPH=1 -e CCL_TOPO_P2P_ACCESS=1 \
  -e CCL_SYCL_ALLREDUCE_SIMPLE_THRESHOLD=1073741824 \
  -e CCL_SYCL_ALLGATHERV_SIMPLE_THRESHOLD=1073741824 \
  -e CCL_SYCL_REDUCE_SCATTER_SIMPLE_THRESHOLD=1073741824 \
  intel/llm-scaler-vllm:0.26.0-b2 /models/qwen38-fp8 --quantization=fp8 \
  -tp 2 --max-model-len 262144 --kv-cache-dtype fp8_e4m3 \
  --max-num-batched-tokens 8192 --block-size 64 --trust-remote-code \
  --max-num-seqs 8 --enable-prefix-caching \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3,"enforce_eager":true}' \
  --mamba-ssm-cache-dtype float16 --gpu-memory-utilization 0.95 --dtype float16 \
  --compilation-config '{"cudagraph_mode":"PIECEWISE","cudagraph_capture_sizes":[4,8,16,32],
    "splitting_ops":["vllm::unified_attention_with_output","vllm::unified_mla_attention_with_output",
    "vllm::mamba_mixer2","vllm::mamba_mixer","vllm::short_conv","vllm::linear_attention",
    "vllm::plamo2_mamba_mixer","vllm::qwen_gdn_attention_core","vllm::gdn_attention_core_xpu",
    "vllm::olmo_hybrid_gdn_full_forward","vllm::kda_attention","vllm::sparse_attn_indexer",
    "vllm::rocm_aiter_sparse_attn_indexer","vllm::deepseek_v4_attention","vllm::hpc_rope_norm_forward",
    "vllm::unified_kv_cache_update","vllm::all_reduce"]}' \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3

Intel's examples run --enforce-eager. Four details of the graph setup that cost me runs: graph mode is on. vllm::all_reduce stays in splitting_ops; when I dropped the splitting ops the all-reduce ended up inside the compiled graph and the output was garbage. The capture sizes are (k+1)×{1,2,4,8} for MTP k=3, where my first guess of 5,10,20,40 was left over from k=4 and padded every decode step by 25%. And the speculative drafter is kept eager.

I also run a small LD_AUDIT shim that stops Level Zero making every allocation resident on the other card. Intel's image does not have it. With the shim and the three oneCCL thresholds together, pinned host RAM went from 24 to 26 GiB per card to 11 GiB per card; I never separated the two, so I can't say how much each contributes. I also wrote a patch to xpu_communicator.py that pads all-reduce messages. It did not stop host-RAM growth on the host-staged path (padded runs still grew 863 to 1,394 MB per worker on the first pass), and only moving to peer-to-peer did. If you run peer-to-peer you probably don't need it. The shim, the patch and the exact stack I tested on are in LocalXPU/b60-dual-fixes (Apache-2.0).

Why not just stay eager

Eager is the safe choice and it never hung on me. All three rows are the same 64 GB host, int4 checkpoint and MTP k=3.

1 stream4 streams8 streams
Graph + MTP73.7 tok/s230361
Compiled, no graphs42.6162313
Eager (--enforce-eager)39.2149288

On real work the gap shrinks but stays. I ran OpenHands on four tasks, one agent at a time, alternating modes overnight. Graph + MTP had a median of 162 s per task against 260 s for eager, and passed 31 of 31 completed trials to eager's 32 of 32. Those tasks can't tell the two apart on quality.

What hung

Starts, on 32 GB. With graphs and MTP, 5 of 19 starts hung while the drafter loaded its weights. The log stopped at Loading safetensors checkpoint shards: 75% | 6/8, and about two seconds later both cards logged this in the same second:

xe 0000:29:00.0: [drm] Tile0: GT0: Engine reset: engine_class=bcs, logical_mask: 0x1, guc_id=32
xe 0000:2d:00.0: [drm] Tile0: GT0: Engine reset: engine_class=bcs, logical_mask: 0x1, guc_id=32

Eager with MTP didn't hang during that load in 15 starts. Stack dumps from three hung starts all sat inside vLLM's weight_loader, in the copy_() of a vocab-sized tensor to the card. Adding a barrier around the drafter load did not stop it: one hang in 5 starts with the first version, one on the 7th start with the second.

Then the RAM. On that host every start pushed available memory down to 0.3 GB with 5 GB of swap in use. On 64 GB I did 45 starts across nine configurations on kernel 6.17, and none hung. I haven't isolated it any further than that, so read it as a strong correlation.

CCL_ALLREDUCE=ring. I tried it after seeing it suggested for TP=2 hangs. On the first concurrent request the host soft-locked for more than 15 minutes, SSH died, and only a reboot cleared it. The trace (trimmed) goes through the xe page-fault worker:

watchdog: BUG: soft lockup - CPU#3 stuck for 652s! [kworker/u49:0]
Workqueue: xe_gt_page_fault_work_queue pf_queue_work_func [xe]
 native_flush_tlb_multi+0x68/0x170
 ttm_bo_unmap_virtual+0x6e/0x80 [ttm]
 xe_bo_validate+0x95/0x130 [xe]
 handle_vma_pagefault+0x139/0x340 [xe]

Don't set it on this setup.

Stalls under load. Graphs with MTP off stalled once in two load runs: four concurrent requests arrived, generation fell to 0 tok/s, and five minutes later the engine died with a timeout in shm_broadcast. On 64 GB, an 8.7 hour overnight agent run alternated graph + MTP with eager and ended with 2.1 hours of deliberate overload on graph + MTP. It logged 0 stalls.

What didn't matter, or made it worse

  • GuC 70.72.1 instead of 70.60.0: 73.4 tok/s against 73.7, 3 starts, no hangs. I put 70.60.0 back.
  • compute-runtime 26.35 with IGC 2.41.5: 69.2 tok/s, 1.2 GiB less VRAM, and the KV cache didn't fit at --gpu-memory-utilization 0.95 (3 starts failed). At 0.90 it ran.
  • llm-scaler 0.26.0-b1: wrong output on the 10-prompt gate (0 of 10), then an xe kernel oops.
  • Mainline kernel 7.2.6 with the stock image: 79 GPU page faults on the first start and resets on both cards. With P2P off there were no new faults, but it hung at the first collective. My guess, untested, is the image's runtime, which was built for 6.17.
  • Mainline 7.3-rc3: 3 of 3 starts failed.

If you do try mainline 7.x on Ubuntu 24.04, the .deb's preinst calls run-parts with two directories and noble's debianutils takes one, so the install stops with "missing operand". I repacked the debs with per-directory calls. Check that the initrd and grub entry actually got made before you reboot; my first repack skipped the hooks and made neither.

Smaller things that cost me time

  • Persist the Triton cache. Without the mount, every start recompiled six kernels (3 to 7 s each) inside the first batched request. That request had 25 s of stalls.
  • The first start after a config change sizes a smaller KV cache. Compiling during vLLM's memory profiling holds up to 2 GiB per card. Mine read 553k, 644k and 637k tokens on cold starts against 657k to 690k warm. Restart once the cache is warm.
  • Graph capture eats host RAM. About 9 GiB per card. Eager used 1.2 GiB per card.
  • Intel's platform eval will show about 122 INT8 TOPS against a 197 spec. I got 121.4 and 122.1 on the two cards at its 30720 size. At 16384 the same card does about 190, in PyTorch and in oneMKL, so I wouldn't go hunting for a fault.

If it hangs anyway

Read the kernel journal, not dmesg. On one of my hosts the dmesg ring buffer held about 1,100 lines, all of them firewall logs, so a reset count from it read zero while the journal had four. journalctl -k --grep 'Engine reset|Timedout job' is the check.

The xe devcoredump is the useful artifact, and it expires. From reading the driver source, xe keeps a snapshot of the first hang only, and while that dump is unread later hangs create nothing, so copy it off and free it. Some resets never produce one: four engine resets over one morning of mine had no coredump line after any of them.

I wrap the start in a watchdog that polls the journal for a new xe reset before vLLM prints "startup complete". If one shows, it captures py-spy stacks of both workers, stops the container and starts again, up to three attempts. A container that exits on its own with no reset is treated as a crash and not retried.

Is Gen3 x8 holding it back?

On the FP8 checkpoint, all-reduce bus bandwidth between the cards measured 4.98 GB/s at large messages and 45 µs at small ones. Prefill ran 1,509 tok/s at 8K, 1,332 at 32K and 1,126 at 64K (58.8 s to first token). In a profile of a 32K prompt, all-reduce took 36.8% of the time. If Gen4 x8 halves the wire time, that is about 22% more prefill, and about 1% more single-stream decode. That is arithmetic on my numbers. I haven't run a Gen4 board.

Still open on my side: a warm FP8 start takes about 300 s, where int4 takes 72 s, and I don't know why.