‹ all posts

Erik Rekola

What 497 cloud agents asked of my desktop

2026-10-08

On the morning of 8 October I checked the fixes my sessions had made overnight. The check ran 497 Claude Sonnet subagents in five workflows. Every model call runs at Anthropic, so none of the reasoning happened on my machine. Commands the agents started did run here, and they kept a 24-core desktop processor busy.

This was a check of my own tooling, not client work. I push limits like this only on my own tools, where I can try whatever I like, and I do not want to spend my own hours checking them. That makes it a good place to see what a run of this size costs the machine it starts from.

What did the agents check?

An outside review on 7 October found faults in my own tools, among them the scripts that send and ship my work. Its findings became a fix plan, and sessions made 70 of those fixes overnight. In the morning each fix got one verifier and two adversaries, whose task was to show the fix was broken. Where anyone claimed a problem, two or three judges ruled on the claim.

Other workflows reran the tools' own tests and looked across the fixes for gaps. Finally, two workflows turned the confirmed findings into a fix list and checked it.

WorkflowAgentsTokensDuration
Each fix: verifier, two adversaries, judges36246 900 1793 h 3 min
Baseline and cross-checks6310 393 9482 h 12 min
Coverage567 691 57352 min
Fix list111 493 43522 min
Finishing5687 61916 min
Total49767 166 754

The verifiers found 62 fixes complete, 8 partial and none missing or harmful. The judges ruled on 388 claims in all. None was high severity. Merging the confirmed claims gave a fix list of 173 items, which the next sessions work through.

Token counts come from the workflow tool itself, about 135 000 per agent. I do not know whether they include input the model read from its cache, so below I treat them as the least an agent reads.

What ran on my machine?

The two largest workflows ran side by side, each with at most 16 agents at a time, so up to 32 agents were running commands at once. They ran the tools' test suites and Go builds of my secrets vault, and browser tests in Playwright's Chromium. In the background the full mutation run of my gate tests went on. It caught all 383 mutations in 2 hours 11 minutes.

I took HWiNFO and Task Manager captures during the run:

ReadingValue
ProcessorIntel Core i9-14900KS, 24 cores and 32 threads
Core usage100 % in one capture, 42 % in another
Package power143 to 146 W, peak 164,7 W
Package temperature55 to 56 °C, peak 65 °C
All-core clockabout 4,6 GHz at about 1,12 V
Memory in use26,1 of 47,8 GB
Processes601, with 8 549 threads and 194 432 handles
RTX 4080 SUPERabout 46 % at 34 to 36 °C

No model ran on the graphics card.

What limited the processor?

HWiNFO names the reason a processor holds its clock back. During the run it named one: IA Electrical Design Point/Other, the current limit often called ICCmax. The cores drew up to 137,5 A. I set that limit conservatively on purpose, and this run was the first time anything reached it. Client work has never come near it.

Power and heat stayed out of it. The power limits PL1 and PL2 stand at 253 W, and the package peaked at 164,7 W. Every thermal reason for the cores stayed at No, and the package peaked at 65 °C.

What would the same check take on a local model?

This is an estimate and not a measurement. The speeds come from other people's published benchmarks, and the rest are my assumptions:

The 4080 SUPER figures come from tests at 32 000 tokens of context, and LocalScore does not state its prompt length. My agents start near 76 000 tokens, and speed falls as the context grows, so the context effect alone makes every time below optimistic.

Machine and modelReadWritePer agent497 agents
RTX 4080 SUPER 16 GB, Qwen3 14B Q41 769 t/s42,6 t/sabout 5 minabout 43 h
RTX 4060 Laptop 8 GB, Llama 3.1 8B Q41 365 t/s36,0 t/sabout 6 minabout 52 h
Core Ultra 9 285H, no discrete graphics card, Llama 3.1 8B Q444 t/s11,3 t/sabout 66 minabout 23 days

The 4080 SUPER figures are hardware-corner.net's RTX 4080 SUPER table at 32k context. The other two rows are LocalScore results for the RTX 4060 Laptop GPU and the Core Ultra 9 285H.

In the cloud the five workflows took 6 hours 45 minutes added together, and the two largest ran side by side.

Would the context even fit?

Not as the agents run today. Qwen3-14B has 40 layers with 8 key and value heads of 128 dimensions each, according to its model configuration. At 16 bits that cache takes 160 KiB per token, so a 76 000 token start needs about 11,6 GiB. Its 4-bit weights take roughly 8 GiB more, about 19,6 GiB in all against a card with 16 GB. Native context for this model is 40 960 tokens, well short of where an agent starts.

Llama 3.1 8B, the model in both laptop rows, takes 128 KiB per token, according to a public copy of its configuration. Its context alone would need about 9,3 GiB, more than the 8 GB of the laptop card. The 14B model would need a quantized cache and a context extension, and the 8B model a quantized cache on the laptop card. I have tested neither.

Why would a laptop not keep up?

The table assumes the machine holds full load for days. My desktop processor drew 143 to 146 W for the tool work alone, with a peak of 164,7 W. Intel rates the Core Ultra 9 285H from the table at 45 W base power and 115 W maximum turbo power. Even that maximum is below what my desktop drew, before a local model takes its own share.

Memory is a second limit. The run had 26,1 GB in use on this machine, more than a laptop with 16 GB has in all. A smaller run would need less, but a local model would also take its share. In my experience a laptop holds a load like this for a few minutes before it slows down. I have not measured that, so the laptop rows are lower bounds and not forecasts.

Would a local model do the same work?

I doubt it, and I have not measured it either. The agents reproduced faults that appear only on Windows and argued against each other's verdicts. Whether an 8B or 14B model can do that is the next thing I would measure. I would run one real verifier prompt on the 4080 and record its speed and duration, then set its verdict next to Sonnet's.

How has the load grown this year?

My own records give a rough series. On 4 July a deep audit ran 4 agents. Audits on 3 and 23 September ran 7 and 10 readers. On 24 September a workspace audit ran 125 agents, and the next evening an audit ran 178. This check ran 497, and I have not yet found where Claude's ceiling for subagents lies.

Part of that jump is a decision of mine. On 29 September I made subagents the default way of working, so the main session plans and delegates and the subagents do the reading and checking. The models got better over the same months, and I have no measure of how much of the growth that explains.

What does this not show?

The local model times are estimates built on other people's benchmarks, and none of them is my own measurement. I did not measure how long a laptop holds this load. I do not know exactly what the workflow tool's token count includes. The machine readings are captures from a few moments, not a log of the whole run, and this is one machine and one run.

Frequently asked

Does a cloud agent load my own computer?

The model runs in the cloud, but the commands it starts run on your machine. In this run up to 32 agents ran tests and builds at once, and the processor package drew 143 to 146 W, with a peak of 164,7 W.

Could a 16 GB graphics card run this check with a local 14B model?

Not as the agents run today. For Qwen3-14B, one agent's starting context of about 76 000 tokens needs about 11,6 GiB of 16-bit cache on top of roughly 8 GiB of 4-bit weights, and the model's native context is 40 960 tokens.

How long would the same check take on a local model?

By my estimate at least 43 hours on an RTX 4080 SUPER, if a quantized cache lets the context fit at all, and about 23 days on a laptop processor without a discrete graphics card, one agent at a time and before the tools' own time. In the cloud the five workflows took 6 hours 45 minutes added together.