Skip to main content

Fireworks AI

info

We support ALL Fireworks AI models, just set fireworks_ai/ as a prefix when sending completion requests

tip

New to running Fireworks AI behind LiteLLM? Getting Started with Fireworks AI on LiteLLM goes from an empty directory to a working request, then adds a second model and a fallback.

PropertyDetails
DescriptionThe fastest and most efficient inference engine to build production-ready, compound AI systems.
Provider Route on LiteLLMfireworks_ai/
Provider DocFireworks AI ↗
Supported OpenAI Endpoints/chat/completions, /responses, /embeddings, /completions, /audio/transcriptions, /rerank

Overview​

This guide explains how to integrate LiteLLM with Fireworks AI. You can connect to Fireworks AI in three main ways:

  1. Using Fireworks AI serverless models – Easy connection to Fireworks-managed models.
  2. Connecting to a model in your own Fireworks account – Access models that are hosted within your Fireworks account.
  3. Connecting via a direct-route deployment – A more flexible, customizable connection to a specific Fireworks instance.

API Key​

# env variable
os.environ['FIREWORKS_AI_API_KEY']

Sample Usage - Serverless Models​

from litellm import completion
import os

os.environ['FIREWORKS_AI_API_KEY'] = ""
response = completion(
model="fireworks_ai/glm-5p3-flash",
messages=[
{"role": "user", "content": "hello from litellm"}
],
)
print(response)

A bare serverless slug like glm-5p3-flash is expanded to accounts/fireworks/models/glm-5p3-flash for you, so you can pass either the short slug or the full resource id.

Sample Usage - Serverless Models - Streaming​

from litellm import completion
import os

os.environ['FIREWORKS_AI_API_KEY'] = ""
response = completion(
model="fireworks_ai/glm-5p3-flash",
messages=[
{"role": "user", "content": "hello from litellm"}
],
stream=True
)

for chunk in response:
print(chunk)

Sample Usage - Models in Your Own Fireworks Account​

from litellm import completion
import os

os.environ['FIREWORKS_AI_API_KEY'] = ""
response = completion(
model="fireworks_ai/accounts/fireworks/models/YOUR_MODEL_ID",
messages=[
{"role": "user", "content": "hello from litellm"}
],
)
print(response)

Sample Usage - Direct-Route Deployment​

from litellm import completion
import os

os.environ['FIREWORKS_AI_API_KEY'] = "YOUR_DIRECT_API_KEY"
response = completion(
model="fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash#accounts/gitlab/deployments/2fb7764c",
messages=[
{"role": "user", "content": "hello from litellm"}
],
api_base="https://gitlab-2fb7764c.direct.fireworks.ai/v1"
)
print(response)

Note: The above is for the chat interface, if you want to use the text completion interface it's model="text-completion-openai/accounts/fireworks/models/deepseek-v4p1-flash#accounts/gitlab/deployments/2fb7764c"

Sample Usage - Routers​

Fireworks routers are served at accounts/fireworks/routers/<router-id> rather than accounts/fireworks/models/<model-id>, so a bare slug alone cannot tell LiteLLM which one you mean. Prefix the slug with routers/ to target a router; LiteLLM expands routers/<id> to accounts/fireworks/routers/<id>. See the Fireworks routers docs for the routers available on your account.

from litellm import completion
import os

os.environ['FIREWORKS_AI_API_KEY'] = ""
response = completion(
model="fireworks_ai/routers/glm-latest",
messages=[
{"role": "user", "content": "hello from litellm"}
],
)
print(response)

The full resource id (fireworks_ai/accounts/fireworks/routers/glm-latest) is still accepted if you prefer to be explicit. Slugs ending in -fast (for example fireworks_ai/glm-5p3-fast) are treated as routers even without the routers/ prefix.

FireRouter and open-model routers​

FireRouter is Fireworks' managed router. Instead of pointing at one model, a router ID picks a model for each user turn. Fireworks serves routers under accounts/fireworks/routers/<id>, and LiteLLM accepts the short ID or the full resource path:

Router IDWhat it routes acrossLiteLLM model
autoFireworks open models, chosen by Fireworksfireworks_ai/auto
auto-instantFireworks open models, tuned for the lowest latencyfireworks_ai/auto-instant
firerouterClaude Opus or GPT, plus the Fireworks open-model mixfireworks_ai/firerouter
firerouter/<models>Only the models you list, such as firerouter/opus or firerouter/kimi-k3/glm-5p3fireworks_ai/firerouter/kimi-k3/glm-5p3
any of the aboveSame router, spelled outfireworks_ai/accounts/fireworks/routers/<id>

auto and auto-instant only use Fireworks open models, so your Fireworks API key is the only credential they need. firerouter/auto and firerouter/auto-instant behave the same way. See Example router IDs for more routes

The full resource path works on every LiteLLM version. The short firerouter IDs need v1.104.0-rc.1 or later, and the short auto and auto-instant IDs need v1.105.0 or later. On older versions a short ID is sent as a model path and Fireworks returns a 404, so use the full path there

model_list:
- model_name: auto
litellm_params:
model: fireworks_ai/accounts/fireworks/routers/auto
api_key: os.environ/FIREWORKS_AI_API_KEY
- model_name: firerouter
litellm_params:
model: fireworks_ai/accounts/fireworks/routers/firerouter
api_key: os.environ/FIREWORKS_AI_API_KEY
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "hello"}]}'

Which model served the request​

The proxy returns the model_name you configured, such as auto, in the response model field. On the SDK, a non-streaming completion() reports the served model in response.model, for example fireworks_ai/glm-5p3-flash. A streamed response keeps the requested router ID in model and puts the served model in response._hidden_params["provider_response_model"]

Closed-model credentials​

Fireworks does not resell closed models. When a firerouter route includes Claude or GPT, Fireworks calls that provider under your own account, using either Provider Keys stored on your Fireworks account or an x-anthropic-api-key or x-openai-api-key header on the request. A header takes precedence over a stored Provider Key

If no credential is available for a closed model, FireRouter leaves it out and serves the turn with the other models in the route. For example, firerouter/opus without an Anthropic credential is served by Fireworks open models. You get 400 no_credential instead when you send x-routing-preference: 1, which forces the route's closed primary, or when you call a closed model ID directly

To send the header from LiteLLM, set it on the deployment with litellm_params.extra_headers, or let each client send its own by enabling forward_client_headers_to_llm_api globally or per model group

model_list:
- model_name: firerouter
litellm_params:
model: fireworks_ai/accounts/fireworks/routers/firerouter
api_key: os.environ/FIREWORKS_AI_API_KEY
extra_headers:
x-anthropic-api-key: os.environ/ANTHROPIC_API_KEY
general_settings:
forward_client_headers_to_llm_api: true
# or per model group:
# model_group_settings:
# forward_client_headers_to_llm_api: [firerouter]

The same headers go through the SDK on completion():

from litellm import completion

response = completion(
model="fireworks_ai/firerouter",
messages=[{"role": "user", "content": "hello"}],
extra_headers={"x-anthropic-api-key": "sk-ant-..."},
)

Routing preference​

x-routing-preference sets how strongly a firerouter request favors its primary model or cheaper models, from 1 (max intelligence) to 5 (max savings). The default is 3. See Routing Preferences. It travels the same way as the credential headers: litellm_params.extra_headers on the deployment, client-supplied when forward_client_headers_to_llm_api is on, or extra_headers on completion()

Cost tracking​

LiteLLM prices each request off the model Fireworks reports it routed to, so a router has no price of its own. Fireworks-hosted models are billed at Fireworks rates, and closed models (for example a Claude turn) are billed at that provider's own list price. The charges land on two vendor bills, Fireworks for open models and Anthropic or OpenAI for closed ones, but LiteLLM spend logs and budgets add both up under the one model group

Usage with LiteLLM Proxy​

1. Set Fireworks AI Models on config.yaml​

model_list:
- model_name: fireworks-glm-5p3
litellm_params:
model: fireworks_ai/glm-5p3-flash
api_key: "os.environ/FIREWORKS_AI_API_KEY"

2. Start Proxy​

litellm --config config.yaml

3. Test it​

curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--data ' {
"model": "fireworks-glm-5p3",
"messages": [
{
"role": "user",
"content": "what llm are you"
}
]
}
'

Responses API​

fireworks_ai/ models on /v1/responses go straight to Fireworks' native https://api.fireworks.ai/inference/v1/responses endpoint, so server-side features such as MCP tools ("type": "mcp"), previous_response_id, and reasoning output items work the same as they do against Fireworks directly

import os
from litellm import responses

os.environ["FIREWORKS_AI_API_KEY"] = "YOUR_API_KEY"

response = responses(
model="fireworks_ai/accounts/fireworks/models/kimi-k3",
input="Use the deepwiki MCP server to tell me in one sentence what the BerriAI/litellm repository is.",
tools=[
{
"type": "mcp",
"server_label": "deepwiki",
"server_url": "https://mcp.deepwiki.com/mcp",
"require_approval": "never",
}
],
)
print(response.output)

Multi-turn tool calling works the same way it does against Fireworks directly: send back the function_call_output items together with the previous_response_id Fireworks returned, and Fireworks continues the conversation server-side

developer input items are sent to Fireworks as system messages, since Fireworks' Responses API has no developer role on models such as kimi-k3 and qwen3.8. A model whose chat template needs the system message first (qwen3.8) still rejects a developer item placed after the first input item, the same way it does when called directly

Document Inlining​

LiteLLM supports document inlining for Fireworks AI models. This is useful for models that are not vision models, but still need to parse documents/images/etc.

LiteLLM will add #transform=inline to the url of the image_url, if the model is not a vision model.See Code

from litellm import completion
import os

os.environ["FIREWORKS_AI_API_KEY"] = "YOUR_API_KEY"
os.environ["FIREWORKS_AI_API_BASE"] = "https://audio-prod.api.fireworks.ai/v1"

completion = litellm.completion(
model="fireworks_ai/accounts/fireworks/models/llama-v3p3-70b-instruct",
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://storage.googleapis.com/fireworks-public/test/sample_resume.pdf"
},
},
{
"type": "text",
"text": "What are the candidate's BA and MBA GPAs?",
},
],
}
],
)
print(completion)

Disable Auto-add​

If you want to disable the auto-add of #transform=inline to the url of the image_url, set disable_add_transform_inline_image_block to True

litellm.disable_add_transform_inline_image_block = True

Reasoning Effort​

The reasoning_effort parameter is supported on select Fireworks AI models. Supported models include:

from litellm import completion
import os

os.environ["FIREWORKS_AI_API_KEY"] = "YOUR_API_KEY"

response = completion(
model="fireworks_ai/accounts/fireworks/models/qwen3-8b",
messages=[
{"role": "user", "content": "What is the capital of France?"}
],
reasoning_effort="low",
)
print(response)

User and Session Attribution​

Fireworks can tell apart the developers behind a shared Fireworks API key when LiteLLM sends two identifiers on each request

IdentifierLiteLLM sourceSent to Fireworks as
User iduser_id of the virtual key that made the requestuser field in the request body
Session idx-litellm-session-id header, litellm_session_id, or metadata.session_idx-session-affinity header

Session id​

fireworks_ai/ deployments send the session id as x-session-affinity on /v1/chat/completions (streaming included), /v1/responses and /v1/messages, so every request in one session reaches the same Fireworks replica and reuses its prompt cache. No configuration is needed

curl http://0.0.0.0:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_KEY" \
-H "x-litellm-session-id: my-session-1" \
-H "Content-Type: application/json" \
-d '{"model": "fireworks-glm-5p2", "messages": [{"role": "user", "content": "hi"}]}'

When the request has no session id, or the proxy generated one because of general_settings.missing_session_id, no x-session-affinity header is sent. A client that sends its own x-session-affinity header keeps that value. For a Fireworks model behind the openai/ prefix, set provider_affinity_header: x-session-affinity in that deployment's litellm_params

User id​

Set fireworks_forward_user_id: true on a deployment to send the LiteLLM user id as Fireworks' user field. It is off by default because user ids are often email addresses

model_list:
- model_name: fireworks-glm-5p2
litellm_params:
model: fireworks_ai/glm-5p2
api_key: "os.environ/FIREWORKS_AI_API_KEY"
fireworks_forward_user_id: true

To turn it on from the Admin UI, open the model on the Models page, click Edit Settings, and add "fireworks_forward_user_id": true to LiteLLM Params

It applies to /v1/chat/completions (streaming included), /v1/responses and /v1/messages on fireworks_ai/ deployments. The value is the user_id of the virtual key, so two developers sharing one Fireworks API key show up as two users, and requests made with the master key send the proxy admin id, default_user_id. Fireworks echoes user back in /v1/responses replies, so the developer sees their own user id there

The LiteLLM user id replaces any user the client sent (and metadata.user_id on /v1/messages) as well as the key hash that litellm_settings.overwrite_user_with_key_hash would send, and a request that tries to set fireworks_forward_user_id itself is rejected, so developers cannot change their own attribution. When the key has no user_id, LiteLLM leaves the request as it is and a client-supplied user is still forwarded. To stop such keys from choosing their own user, also set litellm_settings.overwrite_user_with_key_hash: true, which sends the key hash for them while keys with a user_id still send the user id

Supported Models - ALL Fireworks AI Models Supported!​

info

We support ALL Fireworks AI models, just set fireworks_ai/ as a prefix when sending completion requests

Model NameFunction Call
glm-5p3-flashcompletion(model="fireworks_ai/glm-5p3-flash", messages)
deepseek-v4-procompletion(model="fireworks_ai/deepseek-v4-pro", messages)
kimi-k3completion(model="fireworks_ai/kimi-k3", messages)
qwen3p8-maxcompletion(model="fireworks_ai/qwen3p8-max", messages)
minimax-m3completion(model="fireworks_ai/minimax-m3", messages)
gpt-oss-120bcompletion(model="fireworks_ai/gpt-oss-120b", messages)

The table above is a small selection of popular models. For the full, current list of models and routers, see the Fireworks model library.

Supported Embedding Models​

info

We support ALL Fireworks AI models, just set fireworks_ai/ as a prefix when sending embedding requests

Model NameFunction Call
fireworks_ai/nomic-ai/nomic-embed-text-v1.5response = litellm.embedding(model="fireworks_ai/nomic-ai/nomic-embed-text-v1.5", input=input_text)
fireworks_ai/nomic-ai/nomic-embed-text-v1response = litellm.embedding(model="fireworks_ai/nomic-ai/nomic-embed-text-v1", input=input_text)
fireworks_ai/WhereIsAI/UAE-Large-V1response = litellm.embedding(model="fireworks_ai/WhereIsAI/UAE-Large-V1", input=input_text)
fireworks_ai/thenlper/gte-largeresponse = litellm.embedding(model="fireworks_ai/thenlper/gte-large", input=input_text)
fireworks_ai/thenlper/gte-baseresponse = litellm.embedding(model="fireworks_ai/thenlper/gte-base", input=input_text)

Audio Transcription​

Quick Start​

from litellm import transcription
import os

os.environ["FIREWORKS_AI_API_KEY"] = "YOUR_API_KEY"
os.environ["FIREWORKS_API_BASE"] = "https://audio-prod.api.fireworks.ai/v1"

audio_file = open("/path/to/audio.wav", "rb")

response = transcription(
model="fireworks_ai/whisper-v3",
file=audio_file,
)

Pass API Key/API Base in .transcription

Rerank​

Quick Start​

from litellm import rerank
import os

os.environ["FIREWORKS_AI_API_KEY"] = "YOUR_API_KEY"

query = "What is the capital of France?"
documents = [
"Paris is the capital and largest city of France, home to the Eiffel Tower and the Louvre Museum.",
"France is a country in Western Europe known for its wine, cuisine, and rich history.",
"The weather in Europe varies significantly between northern and southern regions.",
"Python is a popular programming language used for web development and data science.",
]

response = rerank(
model="fireworks_ai/fireworks/qwen3-reranker-8b",
query=query,
documents=documents,
top_n=3,
return_documents=True,
)
print(response)

Pass API Key/API Base in .rerank

Supported Models​

Model NameFunction Call
fireworks/qwen3-reranker-8brerank(model="fireworks_ai/fireworks/qwen3-reranker-8b", query=query, documents=documents)