---
title: "Setting Tag Budgets"
url: "/docs/proxy/tag_budgets"
canonical_url: "https://docs.litellm.ai/docs/proxy/tag_budgets"
type: "docs"
last_updated: "2026-10-09"
summary: "Track spend and set budgets for your API requests using tags. Tags allow you to categorize and monitor costs across different cost centers, projects, and departments."
related:
  - "/docs/proxy/ui_team_soft_budget_alerts"
  - "/docs/proxy/model_access_group_budgets"
---
# Setting Tag Budgets

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


Track spend and set budgets for your API requests using tags. Tags allow you to categorize and monitor costs across different cost centers, projects, and departments.

## Pre-Requisites

- You must set up a Postgres database (e.g. Supabase, Neon, etc.)

## What are Tags?

Tags are labels you can attach to your LLM requests to track and limit spending by category. 

**Common Use Cases:**
- **Cost Center Tracking**: Allocate LLM costs to specific departments or business units (e.g., "engineering", "marketing", "customer-support")
- **Project-based Budgeting**: Set budgets for different projects or initiatives (e.g., "project-alpha", "chatbot-v2")
- **Customer Attribution**: Track spend per customer or client (e.g., "customer-acme", "customer-techcorp")
- **Feature Monitoring**: Monitor costs for specific features (e.g., "feature-chat", "feature-summarization")

Tags can be set on each request (in `metadata` or via `x-litellm-tags`), or attached to a virtual key so every request using that key inherits the tag and its budget limits automatically.

## Setting Tag Budgets

### 1. Create a tag with budget

Create a tag to represent a cost center, project, or any budget category. Set `max_budget` ($ value allowed) and `budget_duration` (how frequently the budget resets).

**Example:** Create a tag for your Engineering department with a monthly $500 budget

#### API

Create a new tag and set `max_budget` and `budget_duration`

```shell
curl -X POST 'http://0.0.0.0:4000/tag/new' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
            "name": "engineering", 
            "description": "Engineering department cost center",
            "max_budget": 500.0, 
            "budget_duration": "30d",
            "rpm_limit": 100,
            "tpm_limit": 100000
        }' 
```

**Request Body Parameters:**

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `name` | string | Yes | Unique name for the tag (e.g., cost center name) |
| `description` | string | No | Description of what this tag tracks |
| `models` | list[string] | No | Restrict tag to specific models |
| `max_budget` | float | No | Maximum budget in USD |
| `budget_duration` | string | No | How often budget resets (e.g., "30d", "1d") |
| `soft_budget` | float | No | Soft budget limit for warnings |
| `rpm_limit` | int | No | Max requests per minute allowed for the tag across all keys and teams |
| `tpm_limit` | int | No | Max tokens per minute allowed for the tag across all keys and teams |

**Response:**

```json
{
  "name": "engineering",
  "description": "Engineering department cost center",
  "max_budget": 500.0,
  "budget_duration": "30d",
  "budget_reset_at": "2025-11-10T00:00:00Z",
  "created_at": "2025-10-11T00:00:00Z"
}  
```

#### LiteLLM Admin UI

Navigate to the **Tag Management** page and click **Create New Tag**. Fill in the tag details and set your budget:

<br />

**Possible values for `budget_duration`:**

| `budget_duration` | When Budget will reset |
| --- | --- |
| `budget_duration="1s"` | every 1 second |
| `budget_duration="1m"` | every 1 minute |
| `budget_duration="1h"` | every 1 hour |
| `budget_duration="1d"` | every 1 day |
| `budget_duration="7d"` | every 1 week |
| `budget_duration="30d"` | every 1 month |

### 2. Attach the tag to an API key (recommended)

Attach the tag when creating or updating a virtual key. Every request made with that key automatically inherits the tag, and the proxy enforces the tag's budget **without** requiring clients to pass `metadata.tags` on each request.

#### API

Use the top-level `tags` field on `/key/generate` or `/key/update`:

```shell
curl -X POST 'http://0.0.0.0:4000/key/generate' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
            "tags": ["engineering"]
        }'
```

You can also set tags under key `metadata`:

```shell
curl -X POST 'http://0.0.0.0:4000/key/generate' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
            "metadata": {
              "tags": ["engineering"]
            }
        }'
```

#### LiteLLM Admin UI

Navigate to **Virtual Keys** → **Create Key** (or edit an existing key) and select the tag(s) in the **Tags** field.

### 3. Use the tag in your requests (optional)

If you did not attach tags to the API key, add tags to each request in the `metadata` field (or via the `x-litellm-tags` header, see [Request Tags](request_tags.md)):

**OpenAI SDK**

```python
import openai

client = openai.OpenAI(
    api_key="sk-<your-litellm-api-key>",  # Your LiteLLM proxy key
    base_url="http://0.0.0.0:4000"
)

response = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={
        "metadata": {
            "tags": ["engineering"]
        }
    }
)
```

**cURL**

```shell
curl -X POST 'http://0.0.0.0:4000/chat/completions' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "model": "gpt-5.6-terra",
           "messages": [{"role": "user", "content": "Hello"}],
           "metadata": {
               "tags": ["engineering"]
           }
         }'
```

### 4. Test It

Make requests with the virtual key from step 2 until the tag budget is exceeded. You do **not** need to pass `metadata.tags` if the tag is already on the key:

```shell
curl -X POST 'http://0.0.0.0:4000/chat/completions' \
     -H 'Authorization: Bearer sk-your-key-with-engineering-tag' \
     -H 'Content-Type: application/json' \
     -d '{
           "model": "gpt-5.6-terra",
           "messages": [{"role": "user", "content": "Hello"}]
         }'
```

If you skipped step 2, include the tag in the request body instead:

```shell
curl -X POST 'http://0.0.0.0:4000/chat/completions' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "model": "gpt-5.6-terra",
           "messages": [{"role": "user", "content": "Hello"}],
           "metadata": {
               "tags": ["engineering"]
           }
         }'
```

**When budget is exceeded, the request is rejected with HTTP 422:**

```json
{
  "error": {
    "message": "Budget has been exceeded! Tag=engineering Current cost: 505.50, Max budget: 500.0",
    "type": "budget_exceeded",
    "param": null,
    "code": "422"
  }
}
```

## Setting Tag Rate Limits

Set `rpm_limit` and `tpm_limit` on a tag to cap requests and tokens per minute for that tag. The limit applies to the tag itself, so usage is shared across every key and team that sends the tag, whether the tag arrives in request `metadata.tags`, the `x-litellm-tags` header, or a key's attached tags. This is separate from the per key `tag_rpm_limit` map in key metadata, which meters each key's requests under a tag independently.

Rate limits on tags are enforced by the v3 parallel request limiter. Once a tag crosses its limit, any request carrying it is rejected with HTTP 429 and a message like `Rate limit exceeded for tag: engineering. Limit type: requests. Current limit: 100`. `rpm_limit` counts each request at admission, and `tpm_limit` is charged with the request's actual token usage after the call completes, so a request can be admitted and the next one rejected once usage lands.

Create a tag with rate limits:

```shell
curl -X POST 'http://0.0.0.0:4000/tag/new' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
            "name": "engineering",
            "rpm_limit": 100,
            "tpm_limit": 100000
        }'
```

Update rate limits on an existing tag:

```shell
curl -X POST 'http://0.0.0.0:4000/tag/update' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
            "name": "engineering",
            "rpm_limit": 200,
            "tpm_limit": 200000
        }'
```

## Managing Tags

### View Tag Information

Get information about specific tags:

```shell
curl -X POST 'http://0.0.0.0:4000/tag/info' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "names": ["engineering", "marketing"]
         }'
```

**Response:**

```json
{
  "engineering": {
    "name": "engineering",
    "description": "Engineering department cost center",
    "spend": 245.50,
    "max_budget": 500.0,
    "budget_duration": "30d",
    "budget_reset_at": "2025-11-10T00:00:00Z",
    "created_at": "2025-10-11T00:00:00Z",
    "updated_at": "2025-10-11T12:30:00Z"
  },
  "marketing": {
    "name": "marketing",
    "description": "Marketing department cost center",
    "spend": 89.20,
    "max_budget": 300.0,
    "budget_duration": "30d",
    "budget_reset_at": "2025-11-10T00:00:00Z",
    "created_at": "2025-10-11T00:00:00Z",
    "updated_at": "2025-10-11T12:30:00Z"
  }
}
```

### Update Tag Budget

Update an existing tag's budget:

```shell
curl -X POST 'http://0.0.0.0:4000/tag/update' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "name": "engineering",
           "max_budget": 750.0,
           "budget_duration": "30d"
         }'
```

### Delete Tag

```shell
curl -X POST 'http://0.0.0.0:4000/tag/delete' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "name": "engineering"
         }'
```

## Multiple Tags per Request

You can apply multiple tags to a single request to track costs across different dimensions simultaneously. For example, track both the cost center and the specific project:

```python
response = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={
        "metadata": {
            "tags": ["engineering", "project-alpha", "customer-acme"]
        }
    }
)
```

```shell
curl -X POST 'http://0.0.0.0:4000/chat/completions' \
     -H "Authorization: Bearer $LITELLM_API_KEY" \
     -H 'Content-Type: application/json' \
     -d '{
           "model": "gpt-5.6-terra",
           "messages": [{"role": "user", "content": "Hello"}],
           "metadata": {
               "tags": ["engineering", "project-alpha", "customer-acme"]
           }
         }'
```

**Budget Enforcement:** If any tag exceeds its budget, the request will be rejected.

## Related pages

- [Team Soft Budget Alerts](https://docs.litellm.ai/docs/proxy/ui_team_soft_budget_alerts.md)
- [Model Access Group Budgets](https://docs.litellm.ai/docs/proxy/model_access_group_budgets.md)
