Two agents and five background workers for two hours
After the 64 GB move I wanted to know if the server would fall over under a messy load. I ran it for about two hours overnight with two OpenHands agents working through tasks, plus five background workers sending prompts of 1k to 30k tokens as fast as the server would take them.
The server took 664 requests, about 7.4 million prompt tokens in all. There were no request errors, no engine stalls, and no GPU error events in the kernel log. Time to first token was 6.8 s at the median and 33.8 s at the 95th percentile. The slow ones are the long prompts waiting behind other long prompts. A cold 44k token prompt takes about 35 s to prefill on its own.
The agents did worse, and it is worth saying why. Six of eight tasks passed. The two that failed were the same ideas task in both rounds, and both hit the 2400 s task limit. The agents were mostly waiting in the queue behind the background prompts, and their output rate fell to about 4 tokens per second against 60 when they have the server to themselves. The server was busy, not broken.
So the server held up. I did not test how it behaves when the queue never drains, and I have not tried more than seven clients.