---
title: "A/B Testing - Traffic Mirroring"
url: "/docs/traffic_mirroring"
canonical_url: "https://docs.litellm.ai/docs/traffic_mirroring"
type: "docs"
last_updated: "2026-10-09"
summary: "Traffic mirroring allows you to \"mimic\" production traffic to a secondary (silent) model for evaluation purposes. The silent model's response is gathered in the background and does not affect the latency or result of the primary request."
related:
  - "/docs/proxy/budget_fallbacks"
  - "/docs/proxy/caching"
---
# A/B Testing - Traffic Mirroring

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


Traffic mirroring allows you to "mimic" production traffic to a secondary (silent) model for evaluation purposes. The silent model's response is gathered in the background and does not affect the latency or result of the primary request.

This is useful for:
- Testing a new model's performance on production prompts before switching.
- Comparing costs and latency between different providers.
- Debugging issues by mirroring traffic to a more verbose model.

## Quick Start

To enable traffic mirroring, add `silent_model` to the `litellm_params` of a deployment.

**SDK**

```python
from litellm import Router

model_list = [
    {
        "model_name": "gpt-5.6-luna",
        "litellm_params": {
            "model": "azure/chatgpt-v-2",
            "api_key": "...",
            "silent_model": "gpt-5.6-terra" # 👈 Mirror traffic to gpt-5.6-terra
        },
    },
    {
        "model_name": "gpt-5.6-terra",
        "litellm_params": {
            "model": "openai/gpt-5.6-terra",
            "api_key": "..."
        },
    }
]

router = Router(model_list=model_list)

# The request to "gpt-5.6-luna" will trigger a background call to "gpt-5.6-terra"
response = await router.acompletion(
    model="gpt-5.6-luna",
    messages=[{"role": "user", "content": "How does traffic mirroring work?"}]
)
```

**Proxy**

Add `silent_model` to your `config.yaml`:

```yaml
model_list:
  - model_name: primary-model
    litellm_params:
      model: azure/gpt-5.6-luna
      api_key: os.environ/AZURE_API_KEY
      silent_model: evaluation-model # 👈 Mirror traffic here
  - model_name: evaluation-model
    litellm_params:
      model: openai/gpt-5.6-terra
      api_key: os.environ/OPENAI_API_KEY
```

## How it works
1. **Request Received**: A request is made to a model group (e.g. `primary-model`).
2. **Deployment Picked**: LiteLLM picks a deployment from the group.
3. **Primary Call**: LiteLLM makes the call to the primary deployment.
4. **Mirroring**: If `silent_model` is present, LiteLLM triggers a background call to that model. 
   - For **Sync** calls: Uses a shared thread pool.
   - For **Async** calls: Uses `asyncio.create_task`.
5. **Isolation**: The background call uses a `deepcopy` of the original request parameters and sets `metadata["is_silent_experiment"] = True`. It also strips out logging IDs to prevent collisions in usage tracking.

## Key Features
- **Latency Isolation**: The primary request returns as soon as it's ready. The background (silent) call does not block.
- **Unified Logging**: Background calls are processed via the Router, meaning they are automatically logged to your configured observability tools (Langfuse, S3, etc.).
- **Evaluation**: Use the `is_silent_experiment: True` flag in your logs to filter and compare results between the primary and mirrored calls.

## Related pages

- [Budget Fallbacks](https://docs.litellm.ai/docs/proxy/budget_fallbacks.md)
- [Caching](https://docs.litellm.ai/docs/proxy/caching.md)
