Graph capture plus MTP against eager mode, on real agent tasks
Decode benchmarks say graph capture plus MTP is about 1.9x faster than eager mode on one stream (73.7 against 39.2 tok/s on this machine). I wanted to know what that looks like when a coding agent does real work, so I ran OpenHands on four tasks and alternated the two setups overnight.
The setup: two Arc Pro B60s, Qwen3.8-27B int4, tensor parallel across both cards, MTP speculative decoding with k=3, Intel’s llm-scaler vLLM image. The only difference between arms is graph capture on or off. The four tasks are an app build, a small game, an ideas task and a business task. I ran six rounds of each setup, alternating, so a slow hour hits both. An earlier run of two trials per task for each setup is counted in the totals below.
| graph + MTP | eager | |
|---|---|---|
| trials passed | 31 of 31 | 32 of 32 |
| median agent time | 162 s | 260 s |
| median wall time | 229 s | 334 s |
| output tokens per agent-second | 59.6 | 38.7 |
That is about 1.6x less agent time (162 against 260 s) and 1.5x less wall time (229 against 334 s). Both are under the 1.9x decode gain because agent time also includes reading the prompt and running tools, and neither of those got faster.
I can’t say anything about quality. Every trial passed in both arms, so these tasks are too easy to tell the two apart. One extra graph trial was thrown out because the task container failed to build and the agent never ran. The server never saw it.
The next comparison needs harder tasks.