---
title: "Amazon Bedrock Mantle"
url: "/docs/providers/bedrock_mantle"
canonical_url: "https://docs.litellm.ai/docs/providers/bedrock_mantle"
type: "docs"
last_updated: "2026-10-09"
summary: "Amazon Bedrock Mantle is Amazon Bedrock's distributed inference engine (Project Mantle) that exposes an OpenAI-compatible API for Bedrock-hosted models."
related:
  - "/docs/providers/bedrock_vector_store"
  - "/docs/providers/litellm_proxy"
---
# Amazon Bedrock Mantle

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


[Amazon Bedrock Mantle](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html) is Amazon Bedrock's distributed inference engine (Project Mantle) that exposes an **OpenAI-compatible API** for Bedrock-hosted models.

Use this provider to call Bedrock Mantle models with accurate **AWS Bedrock pricing** instead of OpenAI pricing.

:::tip

**We support ALL Bedrock Mantle models, just set `model=bedrock_mantle/<model-id>` as a prefix when sending litellm requests**

:::

## Claude Mythos

[Claude Mythos](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-anthropic-claude-mythos-preview.html) (`anthropic.claude-mythos-preview`) is available on Bedrock Mantle with **1M token input context**, 128K output, and support for reasoning, vision, and tool use.

Use the `bedrock_mantle/` route prefix with standard AWS credentials.

### /messages

**SDK**

```python
import asyncio
import litellm
import os

os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-1"

async def main():
    response = await litellm.anthropic_messages(
        model="bedrock_mantle/anthropic.claude-mythos-preview",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
    )
    print(response)

asyncio.run(main())
```

**AI Gateway**

**1. Add to config.yaml**

```yaml
model_list:
  - model_name: claude-mythos
    litellm_params:
      model: bedrock_mantle/anthropic.claude-mythos-preview
      aws_region_name: us-east-1
```

**2. Start LiteLLM AI Gateway**

```shell
litellm --config /path/to/config.yaml
```

**3. Call `/v1/messages` via curl**

```bash
curl -X POST http://0.0.0.0:4000/v1/messages \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -d '{
    "model": "claude-mythos",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Explain quantum entanglement simply."}
    ]
  }'
```

### /chat/completions

**SDK**

```python
from litellm import completion
import os

os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-1"

response = completion(
    model="bedrock_mantle/anthropic.claude-mythos-preview",
    messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
)
print(response)
```

**AI Gateway**

**1. Add to config.yaml**

```yaml
model_list:
  - model_name: claude-mythos
    litellm_params:
      model: bedrock_mantle/anthropic.claude-mythos-preview
      aws_region_name: us-east-1
```

**2. Start LiteLLM AI Gateway**

```shell
litellm --config /path/to/config.yaml
```

**3. Call `/v1/chat/completions` via curl**

```bash
curl -X POST http://0.0.0.0:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -d '{
    "model": "claude-mythos",
    "messages": [
      {"role": "user", "content": "Explain quantum entanglement simply."}
    ]
  }'
```

## Claude Models on /v1/messages

Every `bedrock_mantle/anthropic.claude-*` model, Claude Mythos included, is served on `/v1/messages` from Bedrock Mantle's native Anthropic Messages endpoint, `https://bedrock-mantle.{region}.api.aws/anthropic/v1/messages`, rather than bridged through chat completions, which Mantle rejects for Claude models. This is the surface Claude Code and the Anthropic SDKs talk to, and LiteLLM forwards the request in Anthropic's own wire format, so streaming, tools, and thinking pass straight through. Other Mantle models, the GPT models below for example, keep using the Responses API bridge on `/v1/messages`

Use the bare Mantle model id, such as `bedrock_mantle/anthropic.claude-sonnet-5` or `bedrock_mantle/anthropic.claude-haiku-4-5`. A `us.` inference-profile prefix returns a 404 from Mantle

**SDK**

```python
import asyncio
import litellm
import os

os.environ['AWS_ACCESS_KEY_ID'] = "your-aws-access-key"
os.environ['AWS_SECRET_ACCESS_KEY'] = "your-aws-secret-key"
os.environ['AWS_REGION_NAME'] = "us-east-2"

async def main():
    response = await litellm.anthropic_messages(
        model="bedrock_mantle/anthropic.claude-sonnet-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Explain quantum entanglement simply."}],
    )
    print(response)

asyncio.run(main())
```

**AI Gateway**

**1. Add to config.yaml**

```yaml
model_list:
  - model_name: claude-sonnet-mantle
    litellm_params:
      model: bedrock_mantle/anthropic.claude-sonnet-5
      aws_region_name: us-east-2
      aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
```

**2. Start LiteLLM AI Gateway**

```shell
litellm --config /path/to/config.yaml
```

**3. Call `/v1/messages` via curl**

```bash
curl -X POST http://0.0.0.0:4000/v1/messages \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -d '{
    "model": "claude-sonnet-mantle",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Explain quantum entanglement simply."}
    ]
  }'
```

The region comes from `aws_region_name`, else a region prefix in the model name (`bedrock_mantle/us-east-2/anthropic.claude-sonnet-5`), else the host of `api_base` or `BEDROCK_MANTLE_API_BASE` when it points at Mantle, else `BEDROCK_MANTLE_REGION`, `AWS_REGION_NAME`, or `AWS_REGION`, and finally `us-east-1`. A custom `api_base` (a VPC endpoint or a proxy in front of Mantle) is kept as the host and `/anthropic/v1/messages` is appended to it, whether it was configured with or without an `/openai/v1` or `/v1` suffix

Auth is the same chain as the rest of the provider: a bearer token from `api_key`, `BEDROCK_MANTLE_API_KEY`, or `AWS_BEARER_TOKEN_BEDROCK` when one is set, otherwise SigV4 from `aws_access_key_id` / `aws_secret_access_key` / `aws_session_token`, `aws_profile_name`, or the role params. LiteLLM sends `anthropic-version: 2023-06-01` on every request, and an `anthropic-version` header supplied by the caller wins

Beta features travel in the `anthropic-beta` header: the values the caller sends plus the ones a request needs (a `context_management` edit adds `context-management-2025-06-27`), limited to what Mantle accepts. A value Mantle does not know is left out instead of failing the request with a 400, and nothing is sent in the body `anthropic_beta` field, which Mantle ignores whenever the header is present

Health checks use the same surface. `/health` and the Admin UI's Test Connection button probe a `bedrock_mantle/anthropic.claude-*` deployment with a small `/v1/messages` request, with no `model_info.mode` needed. See [Model modes](../proxy/health.md#model-modes)

## OpenAI Models (GPT-5.4 / GPT-5.5)

### /responses

**SDK**

```python
import litellm
import os

os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"
os.environ['BEDROCK_MANTLE_REGION'] = "us-east-2"

response = litellm.responses(
    model="bedrock_mantle/openai.gpt-5.6-terra",
    input="Hello! How can you help me today?",
)
print(response)
```

#### Streaming

```python
import litellm
import os

os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"

response = litellm.responses(
    model="bedrock_mantle/openai.gpt-5.6-terra",
    input="Tell me a three sentence bedtime story about a unicorn.",
    stream=True,
)

for event in response:
    print(event)
```

**AI Gateway**

**1. Add to config.yaml**

```yaml
model_list:
  - model_name: gpt-5.5-mantle
    litellm_params:
      model: bedrock_mantle/openai.gpt-5.6-terra
      api_key: os.environ/BEDROCK_MANTLE_API_KEY
      api_base: https://bedrock-mantle.us-east-2.api.aws/v1
```

**2. Start LiteLLM AI Gateway**

```shell
litellm --config /path/to/config.yaml
```

**3. Call `/v1/responses` via curl**

```bash
curl -X POST http://0.0.0.0:4000/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -d '{
    "model": "gpt-5.5-mantle",
    "input": "Hello! How can you help me today?"
  }'
```

**4. Or use the OpenAI SDK**

```python
from openai import OpenAI

client = OpenAI(
    api_key="sk-<your-litellm-api-key>",
    base_url="http://0.0.0.0:4000",
)

response = client.responses.create(
    model="gpt-5.5-mantle",
    input="Hello! How can you help me today?",
)
print(response)
```

### Prompt caching (GPT-5.6 and newer)

GPT-5.6 and newer OpenAI models on Mantle accept [explicit prompt cache breakpoints](../completion/prompt_caching.md#bedrock-mantle-explicit-breakpoints-openai-gpt-56-and-newer) on the Responses API: a `prompt_cache_breakpoint` marker on an `input_text`, `input_image` or `input_file` block plus a request-level `prompt_cache_options`. Each cached prefix needs at least 1,024 tokens, a request can carry up to 4 breakpoints, and a cached prefix stays available for at least 30 minutes. LiteLLM passes both fields through on `/v1/responses`. On `/v1/chat/completions` these models are bridged onto the Responses API, so a `prompt_cache_breakpoint` on a content block is kept, and `cache_control_injection_points` on the deployment place the marker for you the way they do for `openai/gpt-5.6` ([tutorial](../tutorials/prompt_caching.md#openai-gpt-56-and-newer))

```yaml
model_list:
  - model_name: gpt-5.6-mantle
    litellm_params:
      model: bedrock_mantle/openai.gpt-5.6-sol
      aws_region_name: us-east-1
      api_key: os.environ/AWS_BEARER_TOKEN_BEDROCK
      cache_control_injection_points:
        - location: message
          role: system
```

```bash
curl -X POST http://0.0.0.0:4000/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-mantle",
    "messages": [
      {"role": "system", "content": "<a system prompt of at least 1,024 tokens>"},
      {"role": "user", "content": "What are the key terms and conditions?"}
    ],
    "prompt_cache_options": {"mode": "explicit", "ttl": "30m"}
  }'
```

The first call reports the cache write in `usage.prompt_tokens_details.cache_write_tokens`, and a repeat of the same prefix within the cache lifetime reports `usage.prompt_tokens_details.cached_tokens`, each priced from the model's cache write and cache read rates

## API Key

```python
# env variable
os.environ['BEDROCK_MANTLE_API_KEY'] = "your-aws-bedrock-api-key"

# optional: override region (defaults to us-east-1)
os.environ['BEDROCK_MANTLE_REGION'] = "us-east-1"  # or use AWS_REGION
```

## Supported Models

| Model | Endpoint | Context Window | Input (per 1M tokens) | Output (per 1M tokens) |
|-------|----------|---------------|----------------------|------------------------|
| `openai.gpt-5.5` | `/responses` | 1.05M | $5.50 | $33.00 |
| `openai.gpt-5.4` | `/responses` | 1.05M | $2.75 | $16.50 |
| `openai.gpt-oss-120b` | `/chat/completions` | 131K | $0.15 | $0.60 |
| `openai.gpt-oss-20b` | `/chat/completions` | 131K | $0.07 | $0.30 |
| `openai.gpt-oss-safeguard-120b` | `/chat/completions` | 131K | $0.15 | $0.60 |
| `openai.gpt-oss-safeguard-20b` | `/chat/completions` | 131K | $0.07 | $0.20 |

## Sample Usage

**SDK**

```python
from litellm import completion
import os

os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"

response = completion(
    model="bedrock_mantle/openai.gpt-oss-120b",
    messages=[{"role": "user", "content": "hello from litellm"}],
)
print(response)
```

**Streaming**

```python
from litellm import completion
import os

os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"

response = completion(
    model="bedrock_mantle/openai.gpt-oss-120b",
    messages=[{"role": "user", "content": "hello from litellm"}],
    stream=True,
)

for chunk in response:
    print(chunk)
```

**Async**

```python
import asyncio
from litellm import acompletion
import os

os.environ['BEDROCK_MANTLE_API_KEY'] = "your-bedrock-api-key"

async def main():
    response = await acompletion(
        model="bedrock_mantle/openai.gpt-oss-120b",
        messages=[{"role": "user", "content": "hello from litellm"}],
    )
    print(response)

asyncio.run(main())
```

## Region Configuration

The API base URL is `https://bedrock-mantle.{region}.api.aws/v1`. Region is resolved in this order:

1. `aws_region_name` on the deployment (or passed as a kwarg)
2. A region prefix in the model name, e.g. `bedrock_mantle/us-gov-west-1/xai.grok-4.3`
3. `BEDROCK_MANTLE_REGION` env var
4. `AWS_REGION_NAME` env var, then `AWS_REGION`
5. Default: `us-east-1`

An explicit `api_base` (or `BEDROCK_MANTLE_API_BASE`) replaces the derived URL entirely. The model-name prefix is stripped before the request is sent, so `bedrock_mantle/us-gov-west-1/xai.grok-4.3` calls `xai.grok-4.3` in `us-gov-west-1`; it is recognized for the regions LiteLLM knows for Bedrock, and `aws_region_name` works for every region. Claude models on `/v1/messages` use the `/anthropic/v1/messages` path instead of `/v1` and keep a custom `api_base` as the host, see [Claude Models on /v1/messages](#claude-models-on-v1messages)

**Supported regions:** `us-east-1`, `us-east-2`, `us-west-2`, `eu-west-1`, `eu-west-2`, `eu-central-1`, `eu-south-1`, `eu-north-1`, `ap-northeast-1`, `ap-south-1`, `ap-southeast-3`, `sa-east-1`, and `us-gov-west-1` (AWS GovCloud)

```python
import os
os.environ['BEDROCK_MANTLE_REGION'] = "eu-west-1"

# or pass api_base directly
response = completion(
    model="bedrock_mantle/openai.gpt-oss-120b",
    messages=[{"role": "user", "content": "hello"}],
    api_base="https://bedrock-mantle.eu-west-1.api.aws/v1",
)
```

### GovCloud pricing

Cost tracking uses the served region. When the price map has a row for `bedrock_mantle/{region}/{model}` (today the `us-gov-west-1` rows), that row prices the call instead of the commercial one, whether the region came from `aws_region_name` or from the model prefix. Both of these deployments bill `xai.grok-4.3` at the GovCloud rate:

```yaml
model_list:
  - model_name: grok-4.3-gov
    litellm_params:
      model: bedrock_mantle/xai.grok-4.3
      aws_region_name: us-gov-west-1
      aws_access_key_id: os.environ/AWS_GOV_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_GOV_SECRET_ACCESS_KEY
  - model_name: grok-4.3-gov-prefixed
    litellm_params:
      model: bedrock_mantle/us-gov-west-1/xai.grok-4.3
      aws_access_key_id: os.environ/AWS_GOV_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_GOV_SECRET_ACCESS_KEY
```

A deployment that sets `base_model` or its own `input_cost_per_token` / `output_cost_per_token` is priced from that row alone; the served region is not applied on top of it

## Usage with LiteLLM Proxy

### 1. Set Bedrock Mantle models on config.yaml

```yaml
model_list:
  - model_name: gpt-5.5-mantle
    litellm_params:
      model: bedrock_mantle/openai.gpt-5.6-terra
      api_key: os.environ/BEDROCK_MANTLE_API_KEY
      api_base: "https://bedrock-mantle.us-east-2.api.aws/v1"

  - model_name: gpt-oss-120b
    litellm_params:
      model: bedrock_mantle/openai.gpt-oss-120b
      api_key: os.environ/BEDROCK_MANTLE_API_KEY
      # optional region override:
      api_base: "https://bedrock-mantle.us-east-1.api.aws/v1"

  - model_name: gpt-oss-20b
    litellm_params:
      model: bedrock_mantle/openai.gpt-oss-20b
      api_key: os.environ/BEDROCK_MANTLE_API_KEY
```

### 2. Start the proxy

```shell
litellm --config /path/to/config.yaml
```

### 3. Send a request

```python
import openai

client = openai.OpenAI(
    api_key="anything",
    base_url="http://0.0.0.0:4000",
)

response = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "hello from litellm"}],
)
print(response)
```

## Related pages

- [Bedrock Knowledge Bases](https://docs.litellm.ai/docs/providers/bedrock_vector_store.md)
- [LiteLLM Proxy (LLM Gateway)](https://docs.litellm.ai/docs/providers/litellm_proxy.md)
