Your router might be costing you more. Not because routing is a bad idea, but because the benchmarks that sell it measure a world your agents do not live in.
RouteLLM’s paper reports cutting costs by over 2x without losing response quality, by sending easy questions to cheap models. That number is real, on the paper’s terms. The paper’s terms are the problem.
The blindspot in the benchmark
RouteLLM’s evaluation is per-query: single-turn MMLU and GSM8K, MT-Bench queries routed independently. Each query is a stranger. No cache to keep warm, nothing lost by switching models, because there is no next turn.
Agent loops are the opposite. Same multi-thousand-token prefix every turn: the system prompt, the tool definitions, the retrieved context, re-sent as one block. And Anthropic prices that block very differently depending on whether it comes from cache: 0.10x for cache reads, 1.25x to write it.
A benchmark with no reuse cannot measure the cost of breaking reuse.
The measurement
We tested the assumption on the Anthropic API on October 6. Same 5,981-token prefix. Run it on one model twice: 5,981 tokens served from cache. Switch to a different model with the identical prefix: 0 read from cache, 5,981 written again at the full 1.25x write price. The cache is per model. Switching models is a cache flush you pay for.
Now multiply that by every turn of every agent loop, and the 2x savings from the paper start running in the wrong direction.
The economics that flip the decision
Here is the part the router vendors do not put in the pitch. Below 1,024 tokens, nothing caches on either model, so routing is genuinely free. Above it, Sonnet serves the cached prefix at 0.10x, which is cheaper than Haiku’s full input price. The cache does your cost optimization for you, better than the router does, as long as you stop switching models.
Pinning the frontier model sounds expensive until you do the arithmetic. Cached Sonnet at 0.10x undercuts Haiku at full price. The “cheap model” is only cheap if you ignore the cache it cannot use.
The decision matrix
Two questions, four cells. Prefix under 1K tokens: route freely, nothing caches down there either way. One turn only: route freely, one turn has nothing to reuse. Prefix over 1K with two or more turns: pin the frontier model, because every switch pays the 1.25x write price again and the 0.10x cached reads are doing your cost optimization already.
Save the matrix before you configure your next agent router. And the next time a routing paper promises 2x, check what the evaluation reuses. If the answer is nothing, the savings are for someone else’s workload.
This post started as an Instagram post →
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.