Amazon Bedrock Mantle
Amazon Bedrock Mantle is Amazon Bedrock's distributed inference engine (Project Mantle) that exposes an OpenAI-compatible API for Bedrock-hosted models.
Use this provider to call Bedrock Mantle models with accurate AWS Bedrock pricing instead of OpenAI pricing.
We support ALL Bedrock Mantle models, just set model=bedrock_mantle/<model-id> as a prefix when sending litellm requests
Claude Mythos
Claude Mythos (anthropic.claude-mythos-preview) is available on Bedrock Mantle with 1M token input context, 128K output, and support for reasoning, vision, and tool use.
Use the bedrock_mantle/ route prefix with standard AWS credentials.
/messages
- SDK
- AI Gateway
import asyncio
import litellm
import os
os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-1"
async def main():
response = await litellm.anthropic_messages(
model="bedrock_mantle/anthropic.claude-mythos-preview",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
)
print(response)
asyncio.run(main())
1. Add to config.yaml
model_list:
- model_name: claude-mythos
litellm_params:
model: bedrock_mantle/anthropic.claude-mythos-preview
aws_region_name: us-east-1
2. Start LiteLLM AI Gateway
litellm --config /path/to/config.yaml
3. Call /v1/messages via curl
curl -X POST http://0.0.0.0:4000/v1/messages \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "claude-mythos",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Explain quantum entanglement simply."}
]
}'
/chat/completions
- SDK
- AI Gateway
from litellm import completion
import os
os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-1"
response = completion(
model="bedrock_mantle/anthropic.claude-mythos-preview",
messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
)
print(response)
1. Add to config.yaml
model_list:
- model_name: claude-mythos
litellm_params:
model: bedrock_mantle/anthropic.claude-mythos-preview
aws_region_name: us-east-1
2. Start LiteLLM AI Gateway
litellm --config /path/to/config.yaml
3. Call /v1/chat/completions via curl
curl -X POST http://0.0.0.0:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "claude-mythos",
"messages": [
{"role": "user", "content": "Explain quantum entanglement simply."}
]
}'
Claude Models on /v1/messages
Every bedrock_mantle/anthropic.claude-* model, Claude Mythos included, is served on /v1/messages from Bedrock Mantle's native Anthropic Messages endpoint, https://bedrock-mantle.{region}.api.aws/anthropic/v1/messages, rather than bridged through chat completions, which Mantle rejects for Claude models. This is the surface Claude Code and the Anthropic SDKs talk to, and LiteLLM forwards the request in Anthropic's own wire format, so streaming, tools, and thinking pass straight through. Other Mantle models, the GPT models below for example, keep using the Responses API bridge on /v1/messages
Use the bare Mantle model id, such as bedrock_mantle/anthropic.claude-sonnet-5 or bedrock_mantle/anthropic.claude-haiku-4-5. A us. inference-profile prefix returns a 404 from Mantle
- SDK
- AI Gateway
import asyncio
import litellm
import os
os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-2"
async def main():
response = await litellm.anthropic_messages(
model="bedrock_mantle/anthropic.claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
)
print(response)
asyncio.run(main())
1. Add to config.yaml
model_list:
- model_name: claude-sonnet-mantle
litellm_params:
model: bedrock_mantle/anthropic.claude-sonnet-5
aws_region_name: us-east-2
aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
2. Start LiteLLM AI Gateway
litellm --config /path/to/config.yaml
3. Call /v1/messages via curl
curl -X POST http://0.0.0.0:4000/v1/messages \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "claude-sonnet-mantle",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Explain quantum entanglement simply."}
]
}'
The region comes from aws_region_name, else a region prefix in the model name (bedrock_mantle/us-east-2/anthropic.claude-sonnet-5), else the host of api_base or BEDROCK_MANTLE_API_BASE when it points at Mantle, else BEDROCK_MANTLE_REGION, AWS_REGION_NAME, or AWS_REGION, and finally us-east-1. A custom api_base (a VPC endpoint or a proxy in front of Mantle) is kept as the host and /anthropic/v1/messages is appended to it, whether it was configured with or without an /openai/v1 or /v1 suffix
Auth is the same chain as the rest of the provider: a bearer token from api_key, BEDROCK_MANTLE_API_KEY, or AWS_BEARER_TOKEN_BEDROCK when one is set, otherwise SigV4 from aws_access_key_id / aws_secret_access_key / aws_session_token, aws_profile_name, or the role params. LiteLLM sends anthropic-version: 2023-06-01 on every request, and an anthropic-version header supplied by the caller wins
Beta features travel in the anthropic-beta header: the values the caller sends plus the ones a request needs (a context_management edit adds context-management-2025-06-27), limited to what Mantle accepts. A value Mantle does not know is left out instead of failing the request with a 400, and nothing is sent in the body anthropic_beta field, which Mantle ignores whenever the header is present
Health checks use the same surface. /health and the Admin UI's Test Connection button probe a bedrock_mantle/anthropic.claude-* deployment with a small /v1/messages request, with no model_info.mode needed. See Model modes
OpenAI Models (GPT-5.4 / GPT-5.5)
/responses
- SDK
- AI Gateway
import litellm
import os
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
os.environ['BEDROCK_MANTLE_REGION'] = "us-east-2"
response = litellm.responses(
model="bedrock_mantle/openai.gpt-5.6-terra",
input="Hello! How can you help me today?",
)
print(response)
Streaming
import litellm
import os
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
response = litellm.responses(
model="bedrock_mantle/openai.gpt-5.6-terra",
input="Tell me a three sentence bedtime story about a unicorn.",
stream=True,
)
for event in response:
print(event)
1. Add to config.yaml
model_list:
- model_name: gpt-5.5-mantle
litellm_params:
model: bedrock_mantle/openai.gpt-5.6-terra
api_key: os.environ/BEDROCK_MANTLE_API_KEY
api_base: https://bedrock-mantle.us-east-2.api.aws/v1
2. Start LiteLLM AI Gateway
litellm --config /path/to/config.yaml
3. Call /v1/responses via curl
curl -X POST http://0.0.0.0:4000/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "gpt-5.5-mantle",
"input": "Hello! How can you help me today?"
}'
4. Or use the OpenAI SDK
from openai import OpenAI
client = OpenAI(
api_key="sk-<your-litellm-api-key>",
base_url="http://0.0.0.0:4000",
)
response = client.responses.create(
model="gpt-5.5-mantle",
input="Hello! How can you help me today?",
)
print(response)
Prompt caching (GPT-5.6 and newer)
GPT-5.6 and newer OpenAI models on Mantle accept explicit prompt cache breakpoints on the Responses API: a prompt_cache_breakpoint marker on an input_text, input_image or input_file block plus a request-level prompt_cache_options. Each cached prefix needs at least 1,024 tokens, a request can carry up to 4 breakpoints, and a cached prefix stays available for at least 30 minutes. LiteLLM passes both fields through on /v1/responses. On /v1/chat/completions these models are bridged onto the Responses API, so a prompt_cache_breakpoint on a content block is kept, and cache_control_injection_points on the deployment place the marker for you the way they do for openai/gpt-5.6 (tutorial)
model_list:
- model_name: gpt-5.6-mantle
litellm_params:
model: bedrock_mantle/openai.gpt-5.6-sol
aws_region_name: us-east-1
api_key: os.environ/AWS_BEARER_TOKEN_BEDROCK
cache_control_injection_points:
- location: message
role: system
curl -X POST http://0.0.0.0:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-mantle",
"messages": [
{"role": "system", "content": "<a system prompt of at least 1,024 tokens>"},
{"role": "user", "content": "What are the key terms and conditions?"}
],
"prompt_cache_options": {"mode": "explicit", "ttl": "30m"}
}'
The first call reports the cache write in usage.prompt_tokens_details.cache_write_tokens, and a repeat of the same prefix within the cache lifetime reports usage.prompt_tokens_details.cached_tokens, each priced from the model's cache write and cache read rates
API Key
# env variable
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-aws-bedrock-api-key"
# optional: override region (defaults to us-east-1)
os.environ['BEDROCK_MANTLE_REGION'] = "us-east-1" # or use AWS_REGION
Supported Models
| Model | Endpoint | Context Window | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|---|---|
openai.gpt-5.5 | /responses | 1.05M | $5.50 | $33.00 |
openai.gpt-5.4 | /responses | 1.05M | $2.75 | $16.50 |
openai.gpt-oss-120b | /chat/completions | 131K | $0.15 | $0.60 |
openai.gpt-oss-20b | /chat/completions | 131K | $0.07 | $0.30 |
openai.gpt-oss-safeguard-120b | /chat/completions | 131K | $0.15 | $0.60 |
openai.gpt-oss-safeguard-20b | /chat/completions | 131K | $0.07 | $0.20 |
Sample Usage
- SDK
- Streaming
- Async
from litellm import completion
import os
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
response = completion(
model="bedrock_mantle/openai.gpt-oss-120b",
messages=[{"role": "user", "content": "hello from litellm"}],
)
print(response)
from litellm import completion
import os
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
response = completion(
model="bedrock_mantle/openai.gpt-oss-120b",
messages=[{"role": "user", "content": "hello from litellm"}],
stream=True,
)
for chunk in response:
print(chunk)
import asyncio
from litellm import acompletion
import os
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
async def main():
response = await acompletion(
model="bedrock_mantle/openai.gpt-oss-120b",
messages=[{"role": "user", "content": "hello from litellm"}],
)
print(response)
asyncio.run(main())
Region Configuration
The API base URL is https://bedrock-mantle.{region}.api.aws/v1. Region is resolved in this order:
aws_region_nameon the deployment (or passed as a kwarg)- A region prefix in the model name, e.g.
bedrock_mantle/us-gov-west-1/xai.grok-4.3 BEDROCK_MANTLE_REGIONenv varAWS_REGION_NAMEenv var, thenAWS_REGION- Default:
us-east-1
An explicit api_base (or BEDROCK_MANTLE_API_BASE) replaces the derived URL entirely. The model-name prefix is stripped before the request is sent, so bedrock_mantle/us-gov-west-1/xai.grok-4.3 calls xai.grok-4.3 in us-gov-west-1; it is recognized for the regions LiteLLM knows for Bedrock, and aws_region_name works for every region. Claude models on /v1/messages use the /anthropic/v1/messages path instead of /v1 and keep a custom api_base as the host, see Claude Models on /v1/messages
Supported regions: us-east-1, us-east-2, us-west-2, eu-west-1, eu-west-2, eu-central-1, eu-south-1, eu-north-1, ap-northeast-1, ap-south-1, ap-southeast-3, sa-east-1, and us-gov-west-1 (AWS GovCloud)
import os
os.environ['BEDROCK_MANTLE_REGION'] = "eu-west-1"
# or pass api_base directly
response = completion(
model="bedrock_mantle/openai.gpt-oss-120b",
messages=[{"role": "user", "content": "hello"}],
api_base="https://bedrock-mantle.eu-west-1.api.aws/v1",
)
GovCloud pricing
Cost tracking uses the served region. When the price map has a row for bedrock_mantle/{region}/{model} (today the us-gov-west-1 rows), that row prices the call instead of the commercial one, whether the region came from aws_region_name or from the model prefix. Both of these deployments bill xai.grok-4.3 at the GovCloud rate:
model_list:
- model_name: grok-4.3-gov
litellm_params:
model: bedrock_mantle/xai.grok-4.3
aws_region_name: us-gov-west-1
aws_access_key_id: os.environ/AWS_GOV_ACCESS_KEY_ID
aws_secret_access_key: os.environ/AWS_GOV_SECRET_ACCESS_KEY
- model_name: grok-4.3-gov-prefixed
litellm_params:
model: bedrock_mantle/us-gov-west-1/xai.grok-4.3
aws_access_key_id: os.environ/AWS_GOV_ACCESS_KEY_ID
aws_secret_access_key: os.environ/AWS_GOV_SECRET_ACCESS_KEY
A deployment that sets base_model or its own input_cost_per_token / output_cost_per_token is priced from that row alone; the served region is not applied on top of it
Usage with LiteLLM Proxy
1. Set Bedrock Mantle models on config.yaml
model_list:
- model_name: gpt-5.5-mantle
litellm_params:
model: bedrock_mantle/openai.gpt-5.6-terra
api_key: os.environ/BEDROCK_MANTLE_API_KEY
api_base: "https://bedrock-mantle.us-east-2.api.aws/v1"
- model_name: gpt-oss-120b
litellm_params:
model: bedrock_mantle/openai.gpt-oss-120b
api_key: os.environ/BEDROCK_MANTLE_API_KEY
# optional region override:
api_base: "https://bedrock-mantle.us-east-1.api.aws/v1"
- model_name: gpt-oss-20b
litellm_params:
model: bedrock_mantle/openai.gpt-oss-20b
api_key: os.environ/BEDROCK_MANTLE_API_KEY
2. Start the proxy
litellm --config /path/to/config.yaml
3. Send a request
import openai
client = openai.OpenAI(
api_key="anything",
base_url="http://0.0.0.0:4000",
)
response = client.chat.completions.create(
model="gpt-oss-120b",
messages=[{"role": "user", "content": "hello from litellm"}],
)
print(response)