Skip to main content

4 posts tagged with "claude-code"

View All Tags

Claude Code server-side auto mode through LiteLLM

LiteLLM Team
LiteLLM Core Team

Last Updated: October 6, 2026

Anthropic is moving Claude Code auto mode's safety classifier from the client to the Claude API. Starting with Claude Code v2.1.278, released September 19, sessions on Enterprise plans and Claude API accounts ask the server to run those checks as part of their own model requests, and Anthropic does not charge for the checks when the server performs them. Anthropic told us the rollout started on September 18 and is gradual, beginning with the Claude Code CLI and VS Code extension and followed by the desktop app and Claude Code on the web over the following week, and that on September 25 auto mode becomes the default permission mode in Claude Code. Today the built-in default is auto on Pro, Max and Team plans and Manual on Enterprise plans and Claude API keys, the accounts that typically sit behind a gateway, per Anthropic's permission modes reference.

Server-side auto mode depends on a contract between Claude Code and the API that some gateways did not preserve, LiteLLM included. The fix first shipped in the dev release cut on Tuesday, September 22, and is now in the v1.104.0 stable release and in patch releases of earlier lines. This post explains what Claude Code needs from an AI Gateway, what LiteLLM was doing wrong, what changed, which release carries it, and how to confirm your deployment is ready.

Incident Report: Prompt Cache Invalidation for Claude Code on Bedrock Invoke

Mateo Wang
AI Engineer, LiteLLM
Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM

Date: July 4 to July 10, 2026
Affected versions: v1.91.0 and v1.91.1
Severity: Medium (silent cost regression; no correctness impact)
Status: Resolved in v1.91.2

Note: If you run Claude Code against Amazon Bedrock through LiteLLM on either v1.91.0 or v1.91.1, upgrade to v1.91.2 or higher.

Summary​

Between July 4 and July 10, proxies running v1.91.0 or v1.91.1 silently broke Anthropic prompt caching for Claude Code sessions routed through Amazon Bedrock's Invoke API. For the customers who reported it, warm-session cache hit rates dropped from roughly 90% to 25-45% and team daily spend rose 2-3x for the same usage. Requests kept returning 200s with correct completions; the only symptoms were the cache miss rate and the bill.

The cause: PR #31364 moved every role: "system" entry in messages into the top-level system field on the Invoke path, which invalidates every cache breakpoint past the tool definitions and system prompt. The fix shipped July 10 in v1.91.2 (#32578, #32831, #32882), with regression tests that fail on pre-fix code.

We own this outcome entirely. The trigger was a poorly documented change in how new Claude models and Claude Code use system messages, but customers run a gateway precisely so they do not have to track provider quirks. Translating requests faithfully, including their caching semantics, is our core job and here we fell short. This post explains exactly what happened, why our testing and review failed to catch it, and what we have changed so this class of regression does not ship again.

5 ways to cut Claude Code costs with LiteLLM

Krrish Dholakia
CEO, LiteLLM

5 ways to save Claude Code cost with LiteLLM

Claude Code is one of the heaviest consumers of input tokens in a modern engineering org. Long tool loops, large file reads, and MCP catalogs with hundreds of tools push every request toward the top of the context window, and the bill scales with it.

If Claude Code already points at a LiteLLM proxy (via ANTHROPIC_BASE_URL), there are five levers the platform admin can pull to bring that cost down. None of them require a client-side change.