Graph capture and MTP finish agent tasks 1.6x faster than eager mode on two Arc Pro B60s

I ran the same 27B model on the same two Arc Pro B60s in three modes and timed an OpenHands agent on each. Graph capture with MTP decodes 73.7 tokens/s on a single stream. Intel's documented eager mode gives 39.2. On the agent tasks the median went from 260 s to 162 s, about 1.6x, and the graphs pin roughly 11 GiB of host RAM per card.

The box is a Ryzen 5 5600 on an MSI X470 Gaming Plus Max with 64 GB of DDR4-3200, both cards at PCIe 3.0 x8, Ubuntu 24.04. Software is Intel's llm-scaler vLLM image 0.26.0-b2, Qwen3.8-27B as an int4 GPTQ checkpoint, tensor parallel across both cards. All three modes run MTP speculative decoding with k=3, so MTP is not the variable. The variable is how vLLM launches the kernels.

EagerCompiledGraph + MTP
1 stream, tok/s39.242.673.7
8 streams, tok/s288313361
Agent time, s260196162
Wall time, s334268229
Trials, all passed32831
Pinned RAM/card, GiB1.22.011.8
Start, s11411572

Eager is --enforce-eager, one launch per operator. Compiled turns on torch.compile but sets the CUDA-graph mode to none. Graph + MTP is what I run: XPU graphs captured at batch sizes 4, 8, 16 and 32, which is (k+1) times 1, 2, 4 and 8.

Theory against measured

Before any of the runs I wrote down what the numbers should be from the model's byte counts and the measured 399 GB/s copy bandwidth per card, with an assumed 40 microseconds per all-reduce. I only modelled the graph + MTP case. For eager and compiled with MTP on there is no theory figure, and I'm not going to make one up afterward.

Graph + MTPExpectedMeasured
1 stream, tok/s80 (65-87)73.7
8 streams, tok/s500 (425-535)361
Pinned RAM/card, GiB0.5-1.511.8 peak
Warm start, s60-19072

Two of those missed. The RAM figure missed because I left graph capture out of the model. A start with 2-second samples showed 3.0 GiB per card climbing to about 10.6 during capture and staying there. Eager measured 1.2, inside the original band, so the extra is the captured graphs. I did not trace which allocation holds it.

The 8-stream decode missed by about 28%. A torch profile of 40 decode steps at 1 and at 8 streams put the step at 39 ms and 63.6 ms. The linear-attention state update kernel went from 1.0 to 7.0 ms where I had budgeted about 3.6, and a gather of the logits across the two cards added 4.1 ms that wasn't in my model at all. That accounts for about 18 of the 24.6 ms of growth. The other 6.6 ms is inside the graph-replayed kernels, which this profiler doesn't record. With those two terms the expectation becomes 360 to 400, and 361 is in it.

Why eager is slow

A decode step on this setup has 138 all-reduces and many small kernels. In eager mode Python launches each one. My reading is that the GPU spends part of every step waiting on the host. I did not profile the eager server, so that is an inference from the size of the gap, not something I observed. Graph capture records the launch sequence once per batch size and replays it, and MTP makes each step produce 2.87 tokens on average at k=3 instead of one.

Compiled without graphs lands next to eager. It decodes about 9% faster at 1 and at 8 streams. On agent tasks its median was 196 s against 260, but that is two trials per task, run after the overnight comparison instead of alternating with it. By task it was slower than eager on three of the four (173 against 155 s, 674 against 536, 261 against 200) and faster on the fourth (143 against 275). One game-logic trial ran 1172 s writing 55k tokens. I would not call it better than eager on this evidence.

Why 1.9x became 1.6x

On one stream graph + MTP decodes 1.9x faster than eager (73.7 against 39.2). On the agent tasks the median agent time is 1.6x lower and the wall time per task 1.5x lower (229 s against 334 s). An agent also reads long prompts and waits on tools, and graph capture doesn't speed those up. At 8 streams the decode gap is only 1.25x (361 against 288), and I haven't isolated why.

The RAM

About 11 GiB per card is 22 GiB of the 64 across the two. On a 32 GB host the same graph + MTP setup hung on 5 of 19 starts, with available memory at 0.3 GB and 5 GB of swap in use on every start. On 64 GB I did 45 starts across all configurations on kernel 6.17 and none hung. I changed other things in between, including a runtime fix (a shim plus oneCCL thresholds) that took pinned RAM from 24 to 26 GiB per card down to about 11, so I can't say how much of that is the RAM alone. The fixes and the config are in the setup guide.

How I measured it

  • Decode. Each mode was started several times (graph + MTP 8, eager 3, compiled 3). After each start a probe sent one stream, then four, then eight concurrent. The table is the median across starts. That is 16 single-stream samples for graph + MTP and 6 each for the others. The compiled number comes from its second run: the first one passed a different setting for the graph mode, produced garbage output (27 tok/s) and is discarded.
  • Agent time. OpenHands 0.62.0 through Harbor 0.23.0, one agent at a time, on four tasks: a CLI feature, game logic, an ideation task and a business report. Agent time is the seconds the agent ran; wall time is the whole trial. Graph + MTP and eager alternated through one night on the same tasks. One extra graph trial was thrown out because the task container failed to build and the agent never ran.
  • Pinned RAM. Peak of the xe driver's system-memory allocation during each soak, per card. Graph mode peaked at 11.8 and 11.0 GiB on the two cards, eager at 1.2 and 0.7, compiled at 2.0 and 1.1. The table shows the larger card.
  • Start time. Seconds from launching the server to it answering. Graph + MTP 81, 73, 72, 72, 73, 72, 72, 72; eager 114, 118, 114; compiled 189, 115, 66.
  • Theory. Bytes read per decode step divided by 399 GB/s per card, plus the all-reduces. The 8-stream expectation was revised after the profile above.

What I can't say

Anything about quality. Every agent trial passed in every mode, so these tasks can't separate them. Graph capture shouldn't change outputs, but I haven't shown that with graded evals. The one deliberate difference from the model's defaults, a float16 recurrent state, matched float32 on 41 of 50 greedy prompts, where two float16 runs match each other on 45 of 50.

The raw per-start summaries and agent timings are not published yet. The next comparison needs harder agent tasks.