Skip to main content

Blog

Insights on routing, reliability, and observability from the team building the most widely used open-source AI gateway.

announcement

Introducing Microsoft 365 Copilot in LiteLLM

Use Microsoft 365 Copilot from your applications through LiteLLM, with access to your Microsoft 365 work data.

Introducing litellm-core: the first step toward a leaner SDK

Ongoing work to reduce LiteLLM's dependencies and installation footprint, improve runtime loading, and make imports faster

anthropic

Claude Haiku 5.5 API Pricing and Day 0 Support on LiteLLM

Claude Haiku 5.5 costs $0.10 per 1M input tokens and $0.50 per 1M output tokens. LiteLLM supports it on day 0 on Anthropic, AWS, Google Cloud and Azure.

agents

Open Sourcing Moyai: Self-Hosted Cloud Coding Agent

Run cloud coding agents on your own infrastructure with Moyai. Choose your agent and models through LiteLLM, track costs, and follow the setup guide.

gemini

Day 0 Support: Nano Banana 2.1

Day 0 support for Gemini Nano Banana 2.1 on LiteLLM, on Google AI Studio and Gemini Enterprise Agent Platform, with image output at half Nano Banana 2's price.

lens

How LiteLLM Lens finds repeated failures across 1,000s of agent traces

Inside Lens: parallel trace review, Python tools for large traces, and investigations that connect repeated failures to source evidence.

performance

How we cut time to first byte by 94% for long prompts

A 440k-token benchmark went from 553 ms to 35 ms median time to first byte. We removed unnecessary prompt-cache routing work before the model call.

engineering

How we built our own internal Devin in 2 days

How we built Moyai Devin with Render, Modal, Hermes, Temporal, and LiteLLM: durable sessions, parallel agents, Slack, and shared organization connections.

auto-router

Adding Self-hosted Auto Router Classifiers: Laya & Nimble

Use Laya or Bespoke Nimble to classify Auto Router requests on your own infrastructure. Control where classification runs, which model you serve, and how you provision it.

lens

Launching LiteLLM Lens

LiteLLM Lens turns the traces flowing through your gateway into findings your agents can act on. Built for agent swarms generating 200K+ traces.

performance

How we cut LiteLLM's Redis round trips per request by 64%

A governed request to the LiteLLM AI Gateway waited on Redis 22 times. It now waits 8 times: one pipeline per Redis backend before the model call, one after.

performance

How we made the LiteLLM Usage page 120x faster

From six minutes to 3.2 seconds: how moving aggregation into Postgres made LiteLLM's Usage page roughly 120x faster in our benchmark.

openai

Day 0 Support: GPT-6.1 Sol

Day 0 support for GPT-6.1 Sol on LiteLLM, with cached input at half GPT-6 Sol's price.