AI agents now burn five times the tokens of human users
Agent traffic on one major AI gateway now dwarfs human traffic five to one, and more than 85% of it is cached re-reading that still has to sit in memory.

Machine agents now consume roughly five times as many tokens as human users on OpenRouter, the model routing platform, and the gap is still widening.
The numbers
Futurum Group chief executive Daniel Newman put the ratio at five to one and predicted it will climb to ten to one and beyond. The chart he shared, drawn from OpenRouter data by the venture firm Andreessen Horowitz, had agents at 7.3 trillion tokens in August against 1.4 trillion for humans, six months after agent traffic first overtook human traffic in February 2026. Since that crossover, OpenRouter's own figures show agent usage up 14-fold while human usage grew 2.8-fold.
Most of it is re-reading
The detail that matters is what the tokens are. More than 85% of agent tokens come from cached prompts, according to a16z, which means the agents are largely re-processing context they have already seen rather than doing new work. A call-centre consultancy's own logs of Claude Code usage in September put it higher still, with 96% of all input being re-reading of earlier conversation.
Cached tokens are cheaper to serve than a fresh prompt, but they still have to be held in memory. Models keep that context in a key-value cache, and the cache is growing faster than the high-bandwidth memory on the GPUs that hold it. One platform's token count is also not a bill, and the trend is not perfectly smooth: OpenRouter's data dips in April and July. The direction, though, is one-way. In McKinsey's 2026 State of AI survey, 40% of respondents at large organisations said they were scaling AI agents, up from 27% a year earlier.
Why it lands on your RAM
The pressure is already visible in the memory market. Micron has said it expects RAM and storage shortages to worsen through 2027 and 2028, with buyers paying more, while memory makers prioritise high-bandwidth memory for AI data centres. Every cached conversation that has to stay resident competes for the same silicon that a graphics card or a desktop upgrade needs, and the agent traffic is not the part that blinks first.
Our opinion
The interesting number here is not five, it is 85. If most agent traffic is a machine re-reading its own notes, then the industry has automated a behaviour it spends a lot of time telling humans off for, and it has done it at a scale where the bill lands on hardware rather than on a subscription line. Caching was supposed to be the optimisation that made agents affordable. Instead it has become the demand signal, because a cache is only worth anything if you keep it in the most expensive memory on the board. The people who feel that first are not the labs. They are anyone who wanted to buy a memory kit this winter.