Skip to main content

Anthropic

LiteLLM supports all anthropic models.

  • claude-haiku-5-5
  • claude-sonnet-5
  • claude-opus-5
  • claude-opus-4-6 (claude-opus-4-6-20260205)
  • claude-sonnet-4-6
  • claude-sonnet-4-5-20250929
  • claude-opus-4-5-20251101
  • claude-opus-4-1-20250805
  • claude-4 (claude-opus-4-20250514, claude-sonnet-4-20250514)
  • claude-3.7 (claude-3-7-sonnet-20250219)
  • claude-3.5 (claude-3-5-sonnet-20240620)
  • claude-3 (claude-3-haiku-20240307, claude-3-opus-20240229, claude-3-sonnet-20240229)
  • claude-2
  • claude-2.1
  • claude-instant-1.2
PropertyDetails
DescriptionClaude is a highly performant, trustworthy, and intelligent AI platform built by Anthropic. Claude excels at tasks involving language, reasoning, analysis, coding, and more. Also available via Azure Foundry.
Provider Route on LiteLLManthropic/ (add this prefix to the model name, to route any requests to Anthropic - e.g. anthropic/claude-3-5-sonnet-20240620). For Azure Foundry deployments, use azure_ai/claude-* (see Azure Anthropic documentation)
Provider DocAnthropic ↗, Azure Foundry Claude ↗
API Endpoint for Providerhttps://api.anthropic.com (or Azure Foundry endpoint: https://<resource-name>.services.ai.azure.com/anthropic)
Supported Endpoints/chat/completions, /v1/messages (passthrough)

Supported OpenAI Parameters​

Check this in code, here

"stream",
"stop",
"temperature",
"top_p",
"max_tokens",
"max_completion_tokens",
"tools",
"tool_choice",
"extra_headers",
"parallel_tool_calls",
"response_format",
"user",
"reasoning_effort",
info

Notes:

  • Anthropic API fails requests when max_tokens are not passed. Due to this litellm passes max_tokens=4096 when no max_tokens are passed.
  • response_format uses Anthropic native structured outputs on Claude Sonnet 4.5+, Opus 4.5+ and Haiku 4.5. Older models such as Opus 4.1 fall back to a forced tool call (see Structured Outputs section)
  • reasoning_effort is automatically mapped to output_config={"effort": ...} for Claude 4.6 and Opus 4.5 models (see Effort Parameter)

Structured Outputs​

LiteLLM supports Anthropic's structured outputs feature for Claude Sonnet 4.5 and later, Opus 4.5 and later, and Haiku 4.5. When you use response_format with these models, LiteLLM automatically:

  • Adds the required structured-outputs-2025-11-13 beta header
  • Transforms OpenAI's response_format to Anthropic's output_format format

Supported Models​

Native structured outputs are used when the model has supports_native_structured_output set in the model cost map:

  • Sonnet 4.5 and later (claude-sonnet-4-5, claude-sonnet-4-6, claude-sonnet-5)
  • Opus 4.5 and later (claude-opus-4-5, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8, claude-opus-5, claude-opus-5-5)
  • Haiku 4.5 and 5.5 (claude-haiku-4-5, claude-haiku-5-5)

Claude Opus 4.1 and older models do not have this flag, so LiteLLM never sends output_format for them. It instead adds a json_tool_call tool built from your schema and forces the model to call it

Example Usage​

from litellm import completion

response = completion(
model="claude-sonnet-5",
messages=[{"role": "user", "content": "What is the capital of France?"}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "capital_response",
"strict": True,
"schema": {
"type": "object",
"properties": {
"country": {"type": "string"},
"capital": {"type": "string"}
},
"required": ["country", "capital"],
"additionalProperties": False
}
}
}
)

print(response.choices[0].message.content)
# Output: {"country": "France", "capital": "Paris"}
info

When using structured outputs with supported models, LiteLLM automatically:

  • Converts OpenAI's response_format to Anthropic's output_format
  • Adds the anthropic-beta: structured-outputs-2025-11-13 header

For models without native support, LiteLLM instead creates a json_tool_call tool with the schema and forces the model to use it

API Keys​

import os

os.environ["ANTHROPIC_API_KEY"] = "your-api-key"
# os.environ["ANTHROPIC_API_BASE"] = "" # [OPTIONAL] or 'ANTHROPIC_BASE_URL'
# os.environ["LITELLM_ANTHROPIC_DISABLE_URL_SUFFIX"] = "true" # [OPTIONAL] Disable automatic URL suffix appending
Azure Foundry Support

Claude models are also available via Microsoft Azure Foundry. Use the azure_ai/ prefix instead of anthropic/ and configure Azure authentication. See the Azure Anthropic documentation for details.

Example:

response = completion(
model="azure_ai/claude-sonnet-5",
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
api_key="your-azure-api-key",
messages=[{"role": "user", "content": "Hello!"}]
)

Custom API Base​

When using a custom API base for Anthropic (e.g., a proxy or custom endpoint), LiteLLM automatically appends the appropriate suffix (/v1/messages or /v1/complete) to your base URL.

If your custom endpoint already includes the full path or doesn't follow Anthropic's standard URL structure, you can disable this automatic suffix appending:

import os

os.environ["ANTHROPIC_API_BASE"] = "https://my-custom-endpoint.com/custom/path"
os.environ["LITELLM_ANTHROPIC_DISABLE_URL_SUFFIX"] = "true" # Prevents automatic suffix

Without LITELLM_ANTHROPIC_DISABLE_URL_SUFFIX:

  • Base URL https://my-proxy.com → https://my-proxy.com/v1/messages
  • Base URL https://my-proxy.com/api → https://my-proxy.com/api/v1/messages

With LITELLM_ANTHROPIC_DISABLE_URL_SUFFIX=true:

  • Base URL https://my-proxy.com/custom/path → https://my-proxy.com/custom/path (unchanged)

Azure AI Foundry (Alternative Method)​

Recommended Method

For full Azure support including Azure AD authentication, use the dedicated Azure Anthropic provider with azure_ai/ prefix.

As an alternative, you can use the anthropic/ provider directly with your Azure endpoint since Azure exposes Claude using Anthropic's native API.

from litellm import completion

response = completion(
model="anthropic/claude-sonnet-5",
api_base="https://<your-resource>.services.ai.azure.com/anthropic",
api_key="<your-azure-api-key>",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response)
info

Finding your Azure endpoint: Go to Azure AI Foundry → Your deployment → Overview. Your base URL will be https://<resource-name>.services.ai.azure.com/anthropic

Workload Identity Federation​

Anthropic supports workload identity federation, so a proxy can exchange an OIDC identity token it already holds for a short-lived sk-ant-oat01 access token instead of storing a long-lived sk-ant- key. This is the Anthropic equivalent of what Vertex AI and Azure already do, and it suits deployments with a no-static-secrets policy.

LiteLLM mints and caches the access token for you. Configure the federation identifiers plus one identity source, and every route that reaches Anthropic uses the minted token: chat completions, /v1/messages, files, batches, skills, passthrough, token counting and model discovery.

Federation identifiers​

These come from the federation rule you create in the Anthropic Console, and they are the same whichever identity source you pick:

FieldDescription
anthropic_federation_rule_idThe fdrl_... id of the federation rule
anthropic_organization_idYour Anthropic organization UUID
anthropic_service_account_idThe svac_... service account the rule maps to
anthropic_federation_workspace_idRequired when the rule is enabled in more than one workspace; scopes the minted token to that workspace

Anthropic's own reference calls the last one workspace_id, but federation reads it from ANTHROPIC_FEDERATION_WORKSPACE_ID, not ANTHROPIC_WORKSPACE_ID. The shorter name already belongs to the Bedrock Claude platform provider, where it picks the workspace a Bedrock call is billed to, so federation takes the longer one and leaves that behavior alone. Setting ANTHROPIC_WORKSPACE_ID does nothing for federation

Static keys take precedence​

A static credential outranks federation everywhere in the Anthropic provider. If ANTHROPIC_API_KEY is set, every Anthropic route uses it and nothing is federated; ANTHROPIC_AUTH_TOKEN comes next, and federation is the last tier. This matches the Anthropic SDK's own ordering, and it applies to files, batches and model discovery as well as to chat.

So a deployment that is meant to be federated must not carry a static key. If one is set anyway, LiteLLM logs a warning naming the model whose federation is being shadowed. A blank or whitespace-only value counts as unset and falls through to federation

Configuring by environment instead​

Every field below can come from the environment rather than the deployment, which is what you want when the same values apply proxy-wide. A value set on the deployment wins over the environment.

FieldEnvironment variable
anthropic_federation_rule_idANTHROPIC_FEDERATION_RULE_ID
anthropic_organization_idANTHROPIC_ORGANIZATION_ID
anthropic_service_account_idANTHROPIC_SERVICE_ACCOUNT_ID
anthropic_federation_workspace_idANTHROPIC_FEDERATION_WORKSPACE_ID
anthropic_identity_sourceANTHROPIC_IDENTITY_SOURCE
anthropic_identity_token_fileANTHROPIC_IDENTITY_TOKEN_FILE
anthropic_identity_tokenANTHROPIC_IDENTITY_TOKEN

Identity sources​

There are four ways to supply the OIDC assertion. Two need no anthropic_identity_source at all, and two are selected with it.

Token file (default)​

Reads an assertion a platform already projects onto disk, which is how Kubernetes service account tokens and most CI runners work. Set anthropic_identity_token_file and leave anthropic_identity_source unset. The path must sit under LiteLLM's OIDC file allowlist, which you extend with LITELLM_OIDC_ALLOWED_CREDENTIAL_DIRS. The file is re-read on each mint, so a rotated token is picked up without a restart.

model_list:
- model_name: claude-sonnet-4-5
litellm_params:
model: anthropic/claude-sonnet-4-5
anthropic_identity_token_file: /var/run/secrets/anthropic.com/token
anthropic_federation_rule_id: os.environ/ANTHROPIC_FEDERATION_RULE_ID
anthropic_organization_id: os.environ/ANTHROPIC_ORGANIZATION_ID
anthropic_service_account_id: os.environ/ANTHROPIC_SERVICE_ACCOUNT_ID

Secret reference (default)​

anthropic_identity_token takes an oidc/ reference, not a raw token. Accepted forms are oidc/env/VAR_NAME, oidc/file//absolute/path, oidc/github/<audience>, and oidc/google/<audience>. Pasting a bare JWT is rejected, so export it and reference the variable instead.

      anthropic_identity_token: oidc/env/MY_WORKLOAD_TOKEN

LiteLLM as the issuer​

For deployments with no external IdP, LiteLLM signs the assertion itself with an ES256 (P-256) key. Set anthropic_identity_source: internal_issuer and point anthropic_issuer_signing_key_ref at the key rather than pasting it inline, so custody stays with your secret manager. anthropic_issuer_ttl_seconds defaults to 300 and cannot exceed 3600.

credential_list:
- credential_name: anthropic_wif
credential_values:
anthropic_identity_source: internal_issuer
anthropic_issuer_url: https://litellm.example.com
anthropic_issuer_subject: litellm-proxy
anthropic_issuer_audience: https://api.anthropic.com
anthropic_issuer_signing_key_ref: os.environ/ANTHROPIC_WIF_SIGNING_KEY
anthropic_issuer_ttl_seconds: 300
anthropic_federation_rule_id: os.environ/ANTHROPIC_FEDERATION_RULE_ID
anthropic_organization_id: os.environ/ANTHROPIC_ORGANIZATION_ID
anthropic_service_account_id: os.environ/ANTHROPIC_SERVICE_ACCOUNT_ID
credential_info:
custom_llm_provider: anthropic

model_list:
- model_name: claude-sonnet-4-5
litellm_params:
model: anthropic/claude-sonnet-4-5
litellm_credential_name: anthropic_wif

Anthropic needs the matching public key to verify what LiteLLM signs. Export it from the proxy and register it as the inline issuer JWKS on your federation rule:

curl -s http://localhost:4000/credentials/anthropic_wif/jwks \
-H "Authorization: Bearer $LITELLM_MASTER_KEY"

The endpoint is proxy-admin only and never exposes the private key, so paste the JSON it returns into the Anthropic Console rather than pointing Anthropic at the URL. A freshly registered JWKS takes about a minute before Anthropic accepts assertions signed by it.

Keycloak​

For shops that already run Keycloak as the workload IdP, LiteLLM fetches the assertion with an OAuth client credentials grant. Set anthropic_identity_source: keycloak plus:

      anthropic_keycloak_token_url: https://keycloak.example.com/realms/prod/protocol/openid-connect/token
anthropic_keycloak_client_id: litellm-proxy
anthropic_keycloak_auth_method: client_secret_basic # or client_secret_post
anthropic_keycloak_client_secret_ref: os.environ/KEYCLOAK_CLIENT_SECRET
anthropic_keycloak_scope: anthropic-federation

Mixing fields across variants is rejected rather than silently ignored, so a config that names internal_issuer while carrying Keycloak fields fails at startup instead of quietly falling back.

Setting it up in the UI​

The Admin UI does not offer these fields yet, so configure federation through the proxy config or the environment as shown above. A follow-up adds them to the LLM Credentials flow.

Token lifetime and refresh​

LiteLLM caches the minted access token per deployment and refreshes it in the background before it expires, so a request rarely waits on an exchange. Refresh is two-tier: an advisory refresh at half the token's lifetime that happens in the background while the old token keeps serving, and a mandatory one at an eighth of the lifetime that blocks. Concurrent requests for the same deployment share a single in-flight exchange rather than each minting their own, so the number of exchanges does not scale with traffic.

Workers on the same host share the minted token through a small on-disk cache under the temp directory, readable only by the proxy's user, so a proxy running several uvicorn workers exchanges each assertion once rather than once per worker. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves that cache, and setting it to an empty string turns it off so every process mints on its own. Replicas on different hosts always mint their own token.

Anthropic accepts each assertion once: a second exchange of the same JWT is denied with the opaque 401 and shows up as jti_reused in the rule's authentication history. That is fine for the internal issuer and Keycloak sources, which mint a fresh assertion for every exchange, but a token file or oidc/env/ value has to rotate faster than the rule's token_lifetime_seconds, or the first refresh after the minted token expires fails until a new assertion shows up. Kubernetes rotates a projected service account token once 80% of its expirationSeconds has passed, so keep the rule's token lifetime at or above that rotation period; the defaults (a one hour projected token and a one hour rule lifetime) line up.

Who can configure it​

Only a proxy admin can create or change a deployment that uses workload identity federation. A team admin who otherwise manages a team-scoped deployment gets a 403, and that holds however the change is expressed: setting a federation field directly, attaching a credential that carries one by name, or changing api_base on a deployment that is already federated. The last one matters because api_base decides where the signed assertion is sent and where the minted token is presented, so it is part of the federation configuration even though it is not a federation field.

The same rule covers the stored fallback lists on a key or team. A fallback target cannot carry a federation field, since those are merged over the deployment's own configuration when a failover happens.

Teams keep normal access to a federated deployment. Only editing it moves to the proxy admin.

Restricting where assertions are sent​

The exchange only talks to api.anthropic.com. If you front Anthropic with a gateway, list its hostname in LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS (comma separated) so the signed assertion is allowed to reach it.

export LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS="anthropic.gateway.internal"

An entry may name a port, in which case only that port is trusted and another process on the same host is not. An entry without a port trusts every port on that host.

export LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS="anthropic.gateway.internal:8443"

The list is read from the environment only. It is never taken from a model or credential API, because api_base decides both where the assertion is sent and where the minted token is presented

Monitoring​

Token health is emitted through the standard service-logging path: prometheus_system sends it to Prometheus and otel sends it to OpenTelemetry, using the exporter settings from the OpenTelemetry docs. The services are anthropic_wif for the exchange itself and anthropic_wif_cache for cache hits and misses, giving you mint counts, mint latency, and failures broken out by cause. Enable them with:

litellm_settings:
service_callback: ["prometheus_system", "otel"]

Usage​

import os
from litellm import completion

# set env - [OPTIONAL] replace with your anthropic key
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

messages = [{"role": "user", "content": "Hey! how's it going?"}]
response = completion(model="claude-opus-5", messages=messages)
print(response)

Usage - Streaming​

Just set stream=True when calling completion.

import os
from litellm import completion

# set env
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

messages = [{"role": "user", "content": "Hey! how's it going?"}]
response = completion(model="claude-opus-5", messages=messages, stream=True)
for chunk in response:
print(chunk["choices"][0]["delta"]["content"]) # same as openai format

Usage with LiteLLM Proxy​

Here's how to call Anthropic with the LiteLLM Proxy Server

1. Save key in your environment​

export ANTHROPIC_API_KEY="your-api-key"

2. Start the proxy​

model_list:
- model_name: claude-4 ### RECEIVED MODEL NAME ###
litellm_params: # all params accepted by litellm.completion() - https://docs.litellm.ai/docs/completion/input
model: claude-opus-5 ### MODEL NAME sent to `litellm.completion()` ###
api_key: "os.environ/ANTHROPIC_API_KEY" # does os.getenv("ANTHROPIC_API_KEY")
litellm --config /path/to/config.yaml

3. Test it​

curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--data ' {
"model": "anthropic/claude-sonnet-5",
"messages": [
{
"role": "user",
"content": "what llm are you"
}
]
}
'

Supported Models​

Model Name 👉 Human-friendly name.
Function Call 👉 How to call the model in LiteLLM.

Model NameFunction Call
claude-opus-4-6completion('claude-opus-4-6-20260205', messages)
claude-sonnet-4-5completion('claude-sonnet-4-5-20250929', messages)
claude-opus-4-5completion('claude-opus-4-5-20251101', messages)
claude-opus-4-1completion('claude-opus-4-1-20250805', messages)
claude-opus-4completion('claude-opus-4-20250514', messages)
claude-sonnet-4completion('claude-sonnet-4-20250514', messages)
claude-3.7completion('claude-3-7-sonnet-20250219', messages)
claude-3-5-sonnetcompletion('claude-3-5-sonnet-20240620', messages)
claude-3-haikucompletion('claude-3-haiku-20240307', messages)
claude-3-opuscompletion('claude-3-opus-20240229', messages)
claude-3-5-sonnet-20240620completion('claude-3-5-sonnet-20240620', messages)
claude-3-sonnetcompletion('claude-3-sonnet-20240229', messages)
claude-2.1completion('claude-2.1', messages)
claude-2completion('claude-2', messages)
claude-instant-1.2completion('claude-instant-1.2', messages)
claude-instant-1completion('claude-instant-1', messages)

Prompt Caching​

Use Anthropic Prompt Caching

Relevant Anthropic API Docs

note

Here's what a sample Raw Request from LiteLLM for Anthropic Context Caching looks like:

POST Request Sent from LiteLLM:
curl -X POST \
https://api.anthropic.com/v1/messages \
-H 'accept: application/json' -H 'anthropic-version: 2023-06-01' -H 'content-type: application/json' -H 'x-api-key: sk-...' \
-d '{'model': 'claude-sonnet-5', [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What are the key terms and conditions in this agreement?",
"cache_control": {
"type": "ephemeral"
}
}
]
},
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "Certainly! The key terms and conditions are the following: the contract is 1 year long for $10/mo"
}
]
}
],
"temperature": 0.2,
"max_tokens": 10
}'

Note: Anthropic no longer requires the anthropic-beta: prompt-caching-2024-07-31 header. Prompt caching now works automatically when you use cache_control in your messages.

Caching - Large Context Caching​

This example demonstrates basic Prompt Caching usage, caching the full text of the legal agreement as a prefix while keeping the user instruction uncached.

response = await litellm.acompletion(
model="anthropic/claude-sonnet-5",
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing legal documents.",
},
{
"type": "text",
"text": "Here is the full text of a complex legal agreement",
"cache_control": {"type": "ephemeral"},
},
],
},
{
"role": "user",
"content": "what are the key terms and conditions in this agreement?",
},
]
)

Caching - Tools definitions​

In this example, we demonstrate caching tool definitions.

The cache_control parameter is placed on the final tool

import litellm

response = await litellm.acompletion(
model="anthropic/claude-sonnet-5",
messages = [{"role": "user", "content": "What's the weather like in Boston today?"}],
tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather in a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
"cache_control": {"type": "ephemeral"}
},
}
]
)

Caching - Continuing Multi-Turn Convo​

In this example, we demonstrate how to use Prompt Caching in a multi-turn conversation.

The cache_control parameter is placed on the system message to designate it as part of the static prefix.

The conversation history (previous messages) is included in the messages array. The final turn is marked with cache-control, for continuing in followups. The second-to-last user message is marked for caching with the cache_control parameter, so that this checkpoint can read from the previous cache.

import litellm

response = await litellm.acompletion(
model="anthropic/claude-sonnet-5",
messages=[
# System Message
{
"role": "system",
"content": [
{
"type": "text",
"text": "Here is the full text of a complex legal agreement"
* 400,
"cache_control": {"type": "ephemeral"},
}
],
},
# marked for caching with the cache_control parameter, so that this checkpoint can read from the previous cache.
{
"role": "user",
"content": [
{
"type": "text",
"text": "What are the key terms and conditions in this agreement?",
"cache_control": {"type": "ephemeral"},
}
],
},
{
"role": "assistant",
"content": "Certainly! the key terms and conditions are the following: the contract is 1 year long for $10/mo",
},
# The final turn is marked with cache-control, for continuing in followups.
{
"role": "user",
"content": [
{
"type": "text",
"text": "What are the key terms and conditions in this agreement?",
"cache_control": {"type": "ephemeral"},
}
],
},
]
)

Function/Tool Calling​

from litellm import completion

# set env
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather in a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
},
}
]
messages = [{"role": "user", "content": "What's the weather like in Boston today?"}]

response = completion(
model="anthropic/claude-sonnet-5",
messages=messages,
tools=tools,
tool_choice="auto",
)
# Add any assertions, here to check response args
print(response)
assert isinstance(response.choices[0].message.tool_calls[0].function.name, str)
assert isinstance(
response.choices[0].message.tool_calls[0].function.arguments, str
)

Forcing Anthropic Tool Use​

If you want Claude to use a specific tool to answer the user’s question

You can do this by specifying the tool in the tool_choice field like so:

response = completion(
model="anthropic/claude-sonnet-5",
messages=messages,
tools=tools,
tool_choice={"type": "tool", "name": "get_weather"},
)

Disable Tool Calling​

You can disable tool calling by setting the tool_choice to "none".

from litellm import completion

response = completion(
model="anthropic/claude-sonnet-5",
messages=messages,
tools=tools,
tool_choice="none",
)

MCP Tool Calling​

Here's how to use MCP tool calling with Anthropic:

LiteLLM supports MCP tool calling with Anthropic in the OpenAI Responses API format.

import os 
from litellm import completion

os.environ["ANTHROPIC_API_KEY"] = "sk-ant-..."

tools=[
{
"type": "mcp",
"server_label": "deepwiki",
"server_url": "https://mcp.deepwiki.com/mcp",
"require_approval": "never",
},
]

response = completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Who won the World Cup in 2022?"}],
tools=tools
)

Parallel Function Calling​

Here's how to pass the result of a function call back to an anthropic model:

from litellm import completion
import os

os.environ["ANTHROPIC_API_KEY"] = "sk-ant.."


litellm.set_verbose = True

### 1ST FUNCTION CALL ###
tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather in a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
},
}
]
messages = [
{
"role": "user",
"content": "What's the weather like in Boston today in Fahrenheit?",
}
]
try:
# test without max tokens
response = completion(
model="anthropic/claude-sonnet-5",
messages=messages,
tools=tools,
tool_choice="auto",
)
# Add any assertions, here to check response args
print(response)
assert isinstance(response.choices[0].message.tool_calls[0].function.name, str)
assert isinstance(
response.choices[0].message.tool_calls[0].function.arguments, str
)

messages.append(
response.choices[0].message.model_dump()
) # Add assistant tool invokes
tool_result = (
'{"location": "Boston", "temperature": "72", "unit": "fahrenheit"}'
)
# Add user submitted tool results in the OpenAI format
messages.append(
{
"tool_call_id": response.choices[0].message.tool_calls[0].id,
"role": "tool",
"name": response.choices[0].message.tool_calls[0].function.name,
"content": tool_result,
}
)
### 2ND FUNCTION CALL ###
# In the second response, Claude should deduce answer from tool results
second_response = completion(
model="anthropic/claude-sonnet-5",
messages=messages,
tools=tools,
tool_choice="auto",
)
print(second_response)
except Exception as e:
print(f"An error occurred - {str(e)}")

s/o @Shekhar Patnaik for requesting this!

Context Management (Beta)​

Anthropic’s context editing API lets you automatically clear older tool results or thinking blocks. LiteLLM now forwards the native context_management payload when you call Anthropic models, and automatically attaches the required context-management-2025-06-27 beta header.

from litellm import completion

response = completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Summarize the latest tool results"}],
context_management={
"edits": [
{
"type": "clear_tool_uses_20250919",
"trigger": {"type": "input_tokens", "value": 30000},
"keep": {"type": "tool_uses", "value": 3},
"clear_at_least": {"type": "input_tokens", "value": 5000},
"exclude_tools": ["web_search"],
}
]
},
)

Anthropic Hosted Tools (Computer, Text Editor, Web Search, Memory)​

from litellm import completion

tools = [
{
"type": "computer_20241022",
"function": {
"name": "computer",
"parameters": {
"display_height_px": 100,
"display_width_px": 100,
"display_number": 1,
},
},
}
]
model = "claude-3-5-sonnet-20241022"
messages = [{"role": "user", "content": "Save a picture of a cat to my desktop."}]

resp = completion(
model=model,
messages=messages,
tools=tools,
# headers={"anthropic-beta": "computer-use-2024-10-22"},
)

print(resp)

Usage - Vision​

from litellm import completion

# set env
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

def encode_image(image_path):
import base64

with open(image_path, "rb") as image_file:
return base64.b64encode(image_file.read()).decode("utf-8")


image_path = "../proxy/cached_logo.jpg"
# Getting the base64 string
base64_image = encode_image(image_path)
resp = litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Whats in this image?"},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64," + base64_image
},
},
],
}
],
)
print(f"\nResponse: {resp}")

Usage - Thinking / reasoning_content​

LiteLLM translates OpenAI's reasoning_effort to Anthropic's thinking parameter. Code

reasoning_effortthinking
"low""budget_tokens": 1024
"medium""budget_tokens": 2048
"high""budget_tokens": 4096
note

reasoning_effort maps to Anthropic's [adaptive thinking](https: //docs.claude.com/en/docs/build-with-claude/extended-thinking/adaptive-thinking) plus the output_config.effort parameter on Claude 4.6 and 4.7 models (including claude-opus-4-6, claude-opus-4-7, claude-sonnet-4-6, etc. ), not budget_tokens. In particular, LiteLLM will inject the following into the underlying Anthropic request on the OpenAI-compatible /chat/completions route:

{
"thinking": {"type": "adaptive"},
"output_config": {"effort": "<low|medium|high|xhigh|max>"}
}

This means any value other than "none" for reasoning_effort will automatically turn thinking on for these models, even though the OpenAI-compatible request body does not have a separate thinking field. This is intended to match Anthropic's own recommended usage: budget_tokens has been deprecated on 4.6 models and rejected entirely on Opus 4.7, where only adaptive is a supported thinking mode.

You can disable thinking either by omitting reasoning_effort entirely or setting it to "none". LiteLLM will not send a thinking field in that case. You can still pass the native thinking parameter directly if you wish to explicitly control thinking with a fixed budget on prior models:

from litellm import completion

# Disable thinking on Claude 4.6/4.7
resp = completion(
model="anthropic/claude-opus-4-7",
messages=[{"role": "user", "content": "What is the capital of France?"}],
reasoning_effort="none", # no thinking field sent
)

# Explicit budget (pre-4.6 models; deprecated on 4.6, rejected on Opus 4.7)
resp = completion(
model="anthropic/claude-sonnet-4-5-20250929",
messages=[{"role": "user", "content": "What is the capital of France?"}],
thinking={"type": "enabled", "budget_tokens": 1024},
)

The Anthropic /v1/messages passthrough route is unaffected by this reasoning effort mapping. thinking is passed through unchanged.

from litellm import completion

resp = completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "What is the capital of France?"}],
reasoning_effort="low",
)

Expected Response

ModelResponse(
id='chatcmpl-c542d76d-f675-4e87-8e5f-05855f5d0f5e',
created=1740470510,
model='claude-sonnet-5',
object='chat.completion',
system_fingerprint=None,
choices=[
Choices(
finish_reason='stop',
index=0,
message=Message(
content="The capital of France is Paris.",
role='assistant',
tool_calls=None,
function_call=None,
provider_specific_fields={
'citations': None,
'thinking_blocks': [
{
'type': 'thinking',
'thinking': 'The capital of France is Paris. This is a very straightforward factual question.',
'signature': 'EuYBCkQYAiJAy6...'
}
]
}
),
thinking_blocks=[
{
'type': 'thinking',
'thinking': 'The capital of France is Paris. This is a very straightforward factual question.',
'signature': 'EuYBCkQYAiJAy6AGB...'
}
],
reasoning_content='The capital of France is Paris. This is a very straightforward factual question.'
)
],
usage=Usage(
completion_tokens=68,
prompt_tokens=42,
total_tokens=110,
completion_tokens_details=None,
prompt_tokens_details=PromptTokensDetailsWrapper(
audio_tokens=None,
cached_tokens=0,
text_tokens=None,
image_tokens=None
),
cache_creation_input_tokens=0,
cache_read_input_tokens=0
)
)

Pass thinking to Anthropic models​

You can also pass the thinking parameter to Anthropic models.

You can also pass the thinking parameter to Anthropic models.

response = litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "What is the capital of France?"}],
thinking={"type": "enabled", "budget_tokens": 1024},
)

Adaptive Thinking (Claude Opus 4.6)​

response = litellm.completion(
model="anthropic/claude-opus-5",
messages=[{"role": "user", "content": "What is the optimal strategy for solving this problem?"}],
thinking={"type": "adaptive"},
)

Enabled Thinking with Budget​

response = litellm.completion(
model="anthropic/claude-opus-5",
messages=[{"role": "user", "content": "What is the capital of France?"}],
thinking={"type": "enabled", "budget_tokens": 5000},
)

Passing Extra Headers to Anthropic API​

Pass extra_headers: dict to litellm.completion

from litellm import completion
messages = [{"role": "user", "content": "What is Anthropic?"}]
response = completion(
model="claude-3-5-sonnet-20240620",
messages=messages,
extra_headers={"anthropic-beta": "max-tokens-3-5-sonnet-2024-07-15"}
)

Usage - "Assistant Pre-fill"​

You can "put words in Claude's mouth" by including an assistant role message as the last item in the messages array.

info

The returned completion will not include your "pre-fill" text, since it is part of the prompt itself. Make sure to prefix Claude's completion with your pre-fill.

import os
from litellm import completion

# set env - [OPTIONAL] replace with your anthropic key
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

messages = [
{"role": "user", "content": "How do you say 'Hello' in German? Return your answer as a JSON object, like this:\n\n{ \"Hello\": \"Hallo\" }"},
{"role": "assistant", "content": "{"},
]
response = completion(model="claude-2.1", messages=messages)
print(response)

Example prompt sent to Claude​


Human: How do you say 'Hello' in German? Return your answer as a JSON object, like this:

{ "Hello": "Hallo" }

Assistant: {

Usage - "System" messages​

If you're using Anthropic's Claude 2.1, system role messages are properly formatted for you.

import os
from litellm import completion

# set env - [OPTIONAL] replace with your anthropic key
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"

messages = [
{"role": "system", "content": "You are a snarky assistant."},
{"role": "user", "content": "How do I boil water?"},
]
response = completion(model="claude-2.1", messages=messages)

Example prompt sent to Claude​

You are a snarky assistant.

Human: How do I boil water?

Assistant:

Mid-conversation system messages, and how LiteLLM places them on each provider so preserved thinking blocks keep their binding, are covered in Preserved Thinking Prefix Stability

Usage - PDF​

Pass base64 encoded PDF files to Anthropic models using the file content type with a file_data field.

using base64​

from litellm import completion, supports_pdf_input
import base64
import requests

# URL of the file
url = "https://storage.googleapis.com/cloud-samples-data/generative-ai/pdf/2403.05530.pdf"

# Download the file
response = requests.get(url)
file_data = response.content

encoded_file = base64.b64encode(file_data).decode("utf-8")

## check if model supports pdf input
supports_pdf_input("anthropic/claude-sonnet-5") # True

response = completion(
model="anthropic/claude-sonnet-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "You are a very professional document summarization specialist. Please summarize the given document."},
{
"type": "file",
"file": {
"file_data": f"data:application/pdf;base64,{encoded_file}", # 👈 PDF
}
},
],
}
],
max_tokens=300,
)

print(response.choices[0])

[BETA] Citations API​

Pass citations: {"enabled": true} to Anthropic, to get citations on your document responses.

Note: This interface is in BETA. If you have feedback on how citations should be returned, please tell us here

from litellm import completion

resp = completion(
model="claude-sonnet-5",
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "text",
"media_type": "text/plain",
"data": "The grass is green. The sky is blue.",
},
"title": "My Document",
"context": "This is a trustworthy document.",
"citations": {"enabled": True},
},
{
"type": "text",
"text": "What color is the grass and sky?",
},
],
}
],
)

citations = resp.choices[0].message.provider_specific_fields["citations"]

assert citations is not None

Files API​

Upload files once and reference them by file_id in multiple requests, with no need to re-upload content each time.

info

The file_id obtained from Anthropic only works with Anthropic Claude models. You cannot use it with other providers (OpenAI, Bedrock, etc.).

  • Max file size: 500 MB | Total storage: 100 GB per org
  • Pricing: File API operations are free. File content used in Messages requests is priced as input tokens.

Supported models by file type:

  • Images: All Claude 3+ models
  • PDFs: All Claude 3.5+ models
  • Other file types (for code execution): Claude 3.5 Haiku + all Claude 3.7+ models

Quick Start​

import litellm
import os

os.environ["ANTHROPIC_API_KEY"] = "sk-ant-..."

# 1. Upload a file once
file = litellm.create_file(
file=open("document.pdf", "rb"),
purpose="messages",
custom_llm_provider="anthropic",
)

# 2. Use file_id in messages (no re-upload needed)
response = litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Summarize this document"},
{"type": "file", "file": {"file_id": file.id, "format": "application/pdf"}}
]
}]
)

File Operations​

OperationFunction
Uploadlitellm.create_file(file, purpose="messages", custom_llm_provider="anthropic")
Listlitellm.file_list(custom_llm_provider="anthropic")
Retrievelitellm.file_retrieve(file_id, custom_llm_provider="anthropic")
Deletelitellm.file_delete(file_id, custom_llm_provider="anthropic")
Downloadlitellm.file_content(file_id, custom_llm_provider="anthropic")
note

Download only works for files created by the code execution tool, not uploaded files.

Supported Formats​

File TypeFormat Value
PDFapplication/pdf
Plain texttext/plain
JPEGimage/jpeg
PNGimage/png
GIFimage/gif
WebPimage/webp

Using Images​

# Upload image
image = litellm.create_file(
file=open("photo.jpg", "rb"),
purpose="messages",
custom_llm_provider="anthropic",
)

# Use in message
response = litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{"type": "file", "file": {"file_id": image.id, "format": "image/jpeg"}}
]
}]
)

Usage - passing 'user_id' to Anthropic​

LiteLLM translates the OpenAI user param to Anthropic's metadata[user_id] param.

response = completion(
model="claude-sonnet-5",
messages=messages,
user="user_123",
)

Usage - Agent Skills​

LiteLLM supports using Agent Skills with the API

response = completion(
model="claude-sonnet-5",
messages=messages,
tools= [
{
"type": "code_execution_20250825",
"name": "code_execution"
}
],
container= {
"skills": [
{
"type": "anthropic",
"skill_id": "pptx",
"version": "latest"
}
]
}
)

The container and its "id" will be present in "provider_specific_fields" in streaming/non-streaming response