Skip to main content

14 posts tagged with "incident-report"

View All Tags

Incident Report: Retry Breadcrumb Memory Growth Causing OOM on v1.100.0

Kerry Lu
Rust Engineer, LiteLLM

Date: September 7 to September 9, 2026
Affected versions: v1.100.0
Severity: Medium (production outage for one customer; low blast radius, required an unreachable Redis)
Status: Resolved in v1.100.1

Note: If you are running v1.100.0, upgrade to v1.100.1 or later. v1.99.1 and earlier are not affected. The v1.100.0 dev and rc pre-releases, v1.101.0-dev.1, and v1.101.0-dev.2 carry the same defect; v1.101.0-rc.1 and later are not affected.

Summary​

Between September 7 and September 9, 2026, an enterprise customer running litellm-database:v1.100.0 behind nginx experienced two waves of kernel OOM kills across their proxy fleet: dozens of kills by the customer's count, one worker process reaching 61 GB RSS, and thousands of HTTP 502s and tens of thousands of HTTP 499s while the platform was down.

This needed two things to be true at once: the provider on the customer's traffic kept failing, so the router kept retrying, and Redis was unreachable at the same time, so every completing request logged the growing result of those retries in full. The retries grew unbounded because of a one-line change in PR #38133 (August 24), which ran retry breadcrumbs, the record the router keeps of each failed LLM attempt, through an existing credential-masking helper before storing them. The helper was correct for its original inputs but assumed it was copying a plain tree; the breadcrumb was not a tree, it contained a pointer back to the same list it was about to be appended to. Copying it without preserving that shared reference turned a structure that had cost nothing in memory since December 2023 into one that roughly doubled in size with every retry, until a cost-tracking error handler turned it into a log string on every request for as long as the Redis outage lasted, and a single allocation reached 27 GB.

We fixed the underlying growth in PR #39491, merged September 3 and included in v1.101.0-rc.1. It was not backported to v1.100.0 before that version was tagged stable two days later. We backported it to stable/1.100.x on September 9 (PR #40455) and published v1.100.1.

We own this outcome entirely. The defect existed in a form we could have caught, a ticket filed three days before the stable release correctly identified a growth problem in the same code path, and our own fix PR turned an untested hedge in that ticket into a stated fact. The rest of this post explains how each of those things happened and what we're changing so that a defect like this can't reach a stable release the same way again.

Incident Report: Prompt Cache Invalidation for Claude Code on Bedrock Invoke

Mateo Wang
AI Engineer, LiteLLM
Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Date: July 4 to July 10, 2026
Affected versions: v1.91.0 and v1.91.1
Severity: Medium (silent cost regression; no correctness impact)
Status: Resolved in v1.91.2

Note: If you run Claude Code against Amazon Bedrock through LiteLLM on either v1.91.0 or v1.91.1, upgrade to v1.91.2 or higher.

Summary​

Between July 4 and July 10, proxies running v1.91.0 or v1.91.1 silently broke Anthropic prompt caching for Claude Code sessions routed through Amazon Bedrock's Invoke API. For the customers who reported it, warm-session cache hit rates dropped from roughly 90% to 25-45% and team daily spend rose 2-3x for the same usage. Requests kept returning 200s with correct completions; the only symptoms were the cache miss rate and the bill.

The cause: PR #31364 moved every role: "system" entry in messages into the top-level system field on the Invoke path, which invalidates every cache breakpoint past the tool definitions and system prompt. The fix shipped July 10 in v1.91.2 (#32578, #32831, #32882), with regression tests that fail on pre-fix code.

We own this outcome entirely. The trigger was a poorly documented change in how new Claude models and Claude Code use system messages, but customers run a gateway precisely so they do not have to track provider quirks. Translating requests faithfully, including their caching semantics, is our core job and here we fell short. This post explains exactly what happened, why our testing and review failed to catch it, and what we have changed so this class of regression does not ship again.

Security Update: Mistral AI PyPI Supply Chain Attack — LiteLLM Not Impacted

On May 11, 2026, security researchers at Aikido Security discovered a coordinated supply chain attack dubbed "Mini Shai-Hulud" that published malicious versions of over 170 npm packages and 2 PyPI packages, including mistralai==2.4.6.

LiteLLM is not impacted. We call Mistral's API directly over HTTP via httpx and do not import the mistralai Python SDK anywhere in the codebase.

TLDR;​

  • LiteLLM does not install or import the mistralai package. We call Mistral's API the same way we call every other provider (via httpx). The compromised package is never executed in any LiteLLM environment.
  • No LiteLLM user credentials were at risk from this attack. The malware runs at import mistralai time. Since LiteLLM never reaches that import, the payload never fires.
  • No action is required from LiteLLM users. If you have separately installed mistralai==2.4.6 in the same environment for your own application code, you should follow Mistral AI's guidance immediately.

Incident Report: Prisma DB Reconnect Blocks the Event Loop and Kills Liveliness

Yuneng Jiang
Senior SWE @ LiteLLM

Date: April 2026 Duration: Multiple incidents across customer deployments before fix landed Severity: High; surfaced as full proxy outages in Kubernetes Status: Resolved

Note: This fix is available starting from the release that contains PR #26225 (merged April 29, 2026).

Summary​

When the upstream Postgres database became unreachable, the LiteLLM proxy's Prisma reconnect path called await self.db.disconnect(). Under prisma-client-py that call invokes a synchronous subprocess.Popen.wait() on the Rust query-engine subprocess. Because wait() does not yield, the asyncio event loop froze for as long as the engine took to shut down, typically 30–120 seconds in production when the engine was stuck on TCP close operations against the unresponsive database.

While the loop was frozen, no coroutines ran, including /health/liveliness. Kubernetes liveness probes timed out and the kubelet SIGKILLed the pod. From the operator's point of view the proxy looked dead even though the underlying issue was a transient DB outage that the reconnect logic was supposed to ride through.

Impact: Any customer whose Postgres briefly became unresponsive saw proxy pods get killed and restarted instead of degrading gracefully and reconnecting once the DB came back. Reported externally by FLock and reproduced internally.

Incident Report: Guardrail logging exposed secret headers in spend logs and traces

LiteLLM Team
LiteLLM Core Team

Date: March 18, 2026 Duration: Unknown Severity: High Status: Resolved

Summary​

When a custom guardrail returned the full LiteLLM request/data dictionary, the guardrail response logged by LiteLLM could include secret_fields.raw_headers, including plaintext Authorization headers containing API keys or other credentials.

This information could then propagate to logging and observability surfaces that consume guardrail metadata, including:

  • Spend logs in the LiteLLM UI: visible to admins with access to spend-log data
  • OpenTelemetry traces: visible to anyone with access to the relevant telemetry backend

LLM calls, proxy routing, and provider execution were not blocked by this bug. The impact was exposure of sensitive request headers in observability and logging paths.

Incident Report: Cache Eviction Closes In-Use httpx Clients

Ryan Crabbe
Performance Engineer, LiteLLM
Ishaan Jaffer
CTO, LiteLLM
Krrish Dholakia
CEO, LiteLLM

Date: February 27, 2026 Duration: ~6 days (Feb 21 merge -> Feb 27 fix) Severity: High Status: Resolved

Note: This fix is available starting from LiteLLM v1.81.14.rc.2 or higher.

Summary​

A change to improve Redis connection pool cleanup introduced a regression that closed httpx clients that were still actively being used by the proxy. The LLMClientCache (an in-memory TTL cache) stores both Redis clients and httpx clients under the same eviction policy. When a cache entry expired or was evicted, the new cleanup code called aclose()/close() on the evicted value which worked correctly for Redis clients, but destroyed httpx clients that other parts of the system still held references to and were actively using for LLM API calls.

Impact: Any proxy instance that hit the cache TTL (default 10 minutes) or capacity limit (200 entries) would have its httpx clients closed out from under it, causing requests to LLM providers to fail with connection errors.

Incident Report: Encrypted Content Failures in Multi-Region Responses API Load Balancing

Sameer Kankute
SWE @ LiteLLM (LLM Translation)
Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Date: Feb 24, 2026
Duration: Ongoing (until fix deployed)
Severity: High (for users load balancing Responses API across different API keys)
Status: Resolved

Summary​

When load balancing OpenAI's Responses API across deployments with different API keys (e.g., different Azure regions or OpenAI organizations), follow-up requests containing encrypted content items (like rs_... reasoning items) would fail with:

{
"error": {
"message": "The encrypted content for item rs_0d09d6e56879e76500699d6feee41c8197bd268aae76141f87 could not be verified. Reason: Encrypted content organization_id did not match the target organization.",
"type": "invalid_request_error",
"code": "invalid_encrypted_content"
}
}

Encrypted content items are cryptographically tied to the API key's organization that created them. When the router load balanced a follow-up request to a deployment with a different API key, decryption failed.

  • Responses API calls with encrypted content: Complete failure when routed to wrong deployment
  • Initial requests: Unaffected. Only follow-up requests containing encrypted items failed
  • Other API endpoints: No impact. Chat completions, embeddings, etc. functioned normally

Incident Report: Wildcard Blocking New Models After Cost Map Reload

Sameer Kankute
SWE @ LiteLLM (LLM Translation)
Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Date: Feb 23, 2026
Duration: ~3 hours
Severity: High (for users with provider wildcard access rules)
Status: Resolved

Summary​

When a new Anthropic model (e.g. claude-sonnet-4-6) was added to the LiteLLM model cost map and a cost map reload was triggered, requests to the new model were rejected with:

key not allowed to access model. This key can only access models=['anthropic/*']. Tried to access claude-sonnet-4-6.

The reload updated litellm.model_cost correctly but never re-ran add_known_models(), so litellm.anthropic_models (the in-memory set used by the wildcard resolver) remained stale. The new model was invisible to the anthropic/* wildcard even though the cost map knew about it.

  • LLM calls: All requests to newly-added Anthropic models were blocked with a 401.
  • Existing models: Unaffected. Only models missing from the stale provider set were impacted.
  • Other providers: Same bug class existed for any provider wildcard (e.g. openai/*, gemini/*).

Incident Report: SERVER_ROOT_PATH regression broke UI routing

Yuneng Jiang
Senior SWE @ LiteLLM
Ishaan Jaffer
CTO, LiteLLM
Krrish Dholakia
CEO, LiteLLM

Date: January 22, 2026 Duration: ~4 days (until fix merged January 26, 2026) Severity: High Status: Resolved

Note: This fix is available starting from LiteLLM v1.81.3.rc.6 or higher.

Summary​

A PR (#19467) accidentally removed the root_path=server_root_path parameter from the FastAPI app initialization in proxy_server.py. This caused the proxy to ignore the SERVER_ROOT_PATH environment variable when serving the UI. Users who deploy LiteLLM behind a reverse proxy with a path prefix (e.g., /api/v1 or /llmproxy) found that all UI pages returned 404 Not Found.

  • LLM API calls: No impact. API routing was unaffected.
  • UI pages: All UI pages returned 404 for deployments using SERVER_ROOT_PATH.
  • Swagger/OpenAPI docs: Broken when accessed through the configured root path.

Incident Report: vLLM Embeddings Broken by encoding_format Parameter

Sameer Kankute
SWE @ LiteLLM (LLM Translation)
Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Date: Feb 16, 2026 Duration: ~3 hours Severity: High (for vLLM embedding users) Status: Resolved

Summary​

A commit (dbcae4a) intended to fix OpenAI SDK behavior broke vLLM embeddings by explicitly passing encoding_format=None in API requests. vLLM rejects this with error: "unknown variant \`, expected float or base64"`.

  • vLLM embedding calls: Complete failure - all requests rejected
  • Other providers: No impact - OpenAI and other providers functioned normally
  • Other vLLM functionality: No impact - only embeddings were affected

Incident Report: Invalid beta headers with Claude Code

Sameer Kankute
SWE @ LiteLLM (LLM Translation)
Ishaan Jaffer
CTO, LiteLLM
Krrish Dholakia
CEO, LiteLLM

Date: February 13, 2026 Duration: ~3 hours Severity: High Status: Resolved

Note: This fix will be available starting from v1.81.13-nightly or higher of LiteLLM.

Summary​

Claude Code began sending unsupported Anthropic beta headers to non-Anthropic providers (Bedrock, Azure AI, Vertex AI), causing invalid beta flag errors. LiteLLM was forwarding all beta headers without provider-specific validation. Users experienced request failures when routing Claude Code requests through LiteLLM to these providers.

  • LLM calls to Anthropic: No impact.
  • LLM calls to Bedrock/Azure/Vertex: Failed with invalid beta flag errors when unsupported headers were present.
  • Cost tracking and routing: No impact.

Incident Report: Invalid model cost map on main

Ishaan Jaffer
CTO, LiteLLM

Date: January 27, 2026 Duration: ~20 minutes Severity: Low Status: Resolved

Summary​

A malformed JSON entry in model_prices_and_context_window.json was merged to main (562f0a0). This caused LiteLLM to silently fall back to a stale local copy of the model cost map. Users on older package versions lost cost tracking for newer models only (e.g. azure/gpt-5.2). No LLM calls were blocked.

  • LLM calls and proxy routing: No impact.
  • Cost tracking: Impacted for newer models not present in the local backup. Older models were unaffected. The incident lasted ~20 minutes until the commit was reverted.