Skip to main content

20 posts tagged with "product"

View All Tags

Introducing the LiteLLM ROI Calculator

Moe Khalil
AI Product Engineer, LiteLLM

LiteLLM ROI Calculator overview: $258 of matched gateway spend compared with 78 estimated engineering hours, or $3.31 per estimated hour, above a list of merged pull requests.

Your gateway tells you what your team spends on AI. It doesn't tell you what that spend produced.

The LiteLLM ROI Calculator compares gateway spend with the engineering work your team ships. It reads spend per user from LiteLLM, estimates the effort in each merged pull request with a model you choose, and matches people by email. The result is one number: spend per estimated engineering hour.

September Townhall Updates: 583 Bug Fixes, OCR on Rust, and 83.5% Coverage

Ishaan Jaffer
CTO, LiteLLM
Yujong Lee
Senior SWE, LiteLLM
Oliver Jensen
Oliver Jensen
Director of Security, LiteLLM
Mateo Wang
AI Engineer, LiteLLM

Thank you to everyone who joined our September town hall. We covered security updates, stability updates, and new features in the product, including OCR running on Rust by default and test coverage past the 80% target we set in August.

TypeSafe Jev on LiteLLM

Kerry Lu
Rust Engineer, LiteLLM

TypeSafe AI's Jev launches on LiteLLM today in v1.103.0-rc. Jev is a decision model: it returns a choice, a score, or a yes/no probability instead of text, so it has its own evaluate endpoint rather than /chat/completions. LiteLLM proxies that endpoint with logging and cost tracking.

Reduce agent context with TypeSafe Jev and LiteLLM

Yassin Kortam
Senior SWE @ LiteLLM

A bot looks up the weather, then checks a shop's opening hours. The user asks, "What time does the shop close?" The bot still sends the old weather report to the model, even though it no longer helps answer the question.

TypeSafe Jev helps LiteLLM spot tool results that are no longer needed. LiteLLM replaces those results with a short notice before calling the model. This is called compaction, and it can reduce the input tokens used by long conversations.

Secure shared AI agents with identity-aware access and spend controls

Yassin Kortam
Senior SWE @ LiteLLM

Shared agents can preserve individual identity, access, and spend controls.

When a finance agent serves multiple business units, platform teams need a consistent way to identify who initiated each request, apply the right model and tool permissions, and attribute spend. LiteLLM keeps this context available across shared-agent workflows so each business unit can operate under its own access and budget policies.

LiteLLM provides one control plane for this workflow across the Agent Gateway, Model Gateway, and MCP Gateway. Teams can share the same agent infrastructure while keeping access, credentials, spend, and audit data tied to the right caller.

Auto Router v1.97: usage benchmarks and better quality for lower cost

Tin Lo
AI Product Engineer, LiteLLM

LiteLLM Autorouter V2: routing accuracy on complex scenarios, 5.6x more accurate by reading the last N turns of the conversation before picking a model



🚀 Help shape the Auto-Router

Get early access, work directly with the LiteLLM team, and influence the roadmap with your production traffic.

Apply to Become a Design Partner

Already testing it? Share your results in discussion #32168.

v1.97 makes three changes to the auto router.

  • The LLM classifier now receives a window of prior conversation turns, defaulting to three. This improves accuracy of follow-up classifications from 14% to 78%, costs at most $0.61 per 1,000 requests, and no additional latency.
  • A new Benchmarks view prices routed traffic against an all-frontier baseline and reports the difference, and those savings now also appear in the Cost Optimization totals.
  • Session affinity is now off by default, following our previous post showing this was leading to worse quality without cost improvements.
Two defaults changed

classifier_context_window_size now defaults to 3 (LLM classifier only), and session_affinity now defaults to false (all routers). Config files are not modified, but the new defaults apply to any key left unset, so a config that never mentioned session_affinity will reclassify every turn after upgrading. Configs that set either key explicitly are unaffected.

Announcing Router Plugins: Customize Routing Signals

Krrish Dholakia
CEO, LiteLLM
Availability

Router plugins run on the proxy from v1.94.x. The design is still evolving; tell us how you'd use it and what you'd want next in the autorouter discussion on GitHub (#32168).

Router plugins are now available on LiteLLM. Each plugin receives the routing context, enriches it, and hands it to the next before the router makes the final decision.

The push came from the autorouter discussion (#32168): teams wanted to layer their own signals (language detection, domain classification, tenant policy, budget caps) onto routing without waiting for each one to land in core. This plugin extension lets teams make these changes while keeping LiteLLM's routing core stable.

Auto Router v2: one router for complexity, semantic, and adaptive routing

Krrish Dholakia
CEO, LiteLLM
Availability

Auto Router v2 ships in v1.94.x. The earliest dev release cuts Tuesday, 2026-07-14. Suggestions and feedback: discussion #32168.

Auto Router v2 collapses complexity, semantic, and adaptive routing into a single auto_router/complexity_router. One config now covers heuristic scoring, LLM classification, lexical or semantic keyword rules, and Thompson-sampled tier pools.

The push came from the community. On discussion #32168, users pointed out that all three routing strategies should converge into a single Auto Router. One router with configurable signals and weights keeps the API simple while letting the routing engine evolve internally, instead of forcing you to pick a mode up front.

The operational half came from discussion #32172: predictable beats clever for debuggability. A fixed, versioned mapping from capability class to model is what makes "why did this response cost 4x today" answerable after the fact.

July stability update: hardening MCP auth and cutting pass-through memory

Ishaan Jaffer
CTO, LiteLLM
Tin Lo
AI Product Engineer, LiteLLM
Mateo Wang
AI Engineer, LiteLLM
Yassin Kortam
Senior SWE @ LiteLLM

Over the last two weeks we addressed two major product quality issues:

  1. The MCP Gateway did not have a single class for credential resolution.
  2. Pass-through APIs had high memory consumption.

Across the same window we shipped 134 bug fixes in total. This post covers the two big changes first, then the rest of the AI Eng and reliability work, the full breakdown, and what we are doing next.

June Townhall Updates: 94 Bug Fixes, OCR + Realtime are in Rust, and a Zero-Regression Commitment

Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Thank you to everyone who joined our June town hall.

Three numbers capture the month: 24 security fixes, 94 bug fixes, and 78 feature commits. The sections below break each one down, alongside our public commitment to zero reported regressions and the gradual migration of the LiteLLM gateway to Rust.

LiteLLM Labs: Announcing Lite-Harness SDK — Unified API for Claude Code, Codex, and Pi AI

Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Harnesses are the next frontier of vendor lock-in. LiteLLM was built to swap across model providers easily. However, as the models get saturated, the next area for competition becomes the harnesses and managed agents. To make it easy to go across vendors at the harness layer, we're launching the Lite-Harness SDK. This is a simple TypeScript+Python SDK which allows developers to change harnesses, like they change models.