Auto Router: Maximize Quality & Savings with Our Fuse LLM Classifier
Weβre improving Fuse V2 by combining complexity and capability assessment to learn where each model succeeds, fails, and needs an upgrade.
Weβre improving Fuse V2 by combining complexity and capability assessment to learn where each model succeeds, fails, and needs an upgrade.
We solved 23 of 25 SWE-bench Verified tasks with LiteLLM's experimental capability router for $11.15. With Opus 5 for all solver calls, we solved 23 tasks for $20.27. We spent 45% less, including classifier calls, with prompt caching on in both runs.

An experimental router that picks a model per agent phase matched fixed Claude Opus-5 quality on a SWE-bench Verified subset for 46% less money.

The complexity router's LLM classifier can now be compressed more aggressively than your model calls. In internal testing, that cut classification costs a further 32% beyond what shared compression was already saving, with no change in routing accuracy.

Heuristic v2, LiteLLM's new AutoRouter classifier, is up to 45% more efficient than Heuristic v1: more tasks solved, at lower cost, in less time. Only the classifier changed.

LiteLLM Auto Router Fusion solved 14 of 21 Terminal-Bench tasks; Claude Fable-5 on its own solved 9. Fusion runs the task on several models in parallel and has one of them synthesize the candidate work into a single answer. Both arms ran the same 21 tasks.
You can expect roughly 40% cost reductions from day one with the Auto Router, and more as the tier maps are tuned. One of our production users shared their statistics to show what that looks like at scale.
They rolled out the Auto-Router to 450+ users across dev, staging, and prod instances and saved $12,249 over 270k+ requests.


An auto router matched Claude Opus-5 solve rate on a 21 task subset of Terminal-Bench 2.0 at 27% lower cost. Every arm ran the same 21 tasks, so the comparisons are like for like.

Get early access, work directly with the LiteLLM team, and influence the roadmap with your production traffic.
Apply to Become a Design PartnerAlready testing it? Share your results in discussion #32168.
v1.97 makes three changes to the auto router.
classifier_context_window_size now defaults to 3 (LLM classifier only), and session_affinity now defaults to false (all routers). Config files are not modified, but the new defaults apply to any key left unset, so a config that never mentioned session_affinity will reclassify every turn after upgrading. Configs that set either key explicitly are unaffected.

Yes, you can use prompt caching with Auto-Routing. The two compound rather than cancel out. We measured it across five datasets, two of which report what the provider's cache actually did.

Auto routing promises a smaller bill without a worse answer. We measured both halves against a baseline that sends every request to claude-opus-5: 8,619 graded prompts and cost simulations over 14,000 real conversations.

Claude Code is one of the heaviest consumers of input tokens in a modern engineering org. Long tool loops, large file reads, and MCP catalogs with hundreds of tools push every request toward the top of the context window, and the bill scales with it.
If Claude Code already points at a LiteLLM proxy (via ANTHROPIC_BASE_URL), there are five levers the platform admin can pull to bring that cost down. None of them require a client-side change.