Overview
| Feature | Supported |
|---|---|
| Supported Providers | perplexity, tavily, parallel_ai, exa_ai, brave, google_pse, dataforseo, firecrawl, searxng, linkup, duckduckgo, searchapi, serper, you_com, apiserpent, agentcore, nimble, bing_grounding |
| Cost Tracking | ✅ |
| Logging | ✅ |
| Load Balancing | ❌ |
LiteLLM follows the Perplexity API request/response for the Search API
Supported from LiteLLM v1.78.7+
LiteLLM Python SDK Usage
Quick Start
from litellm import search
import os
os.environ["PERPLEXITYAI_API_KEY"] = "pplx-..."
response = search(
query="latest AI developments in 2024",
search_provider="perplexity",
max_results=5
)
# Access search results
for result in response.results:
print(f"{result.title}: {result.url}")
print(f"Snippet: {result.snippet}\n")
To use Parallel AI Search, set PARALLEL_API_KEY and pass search_provider="parallel_ai".
Async Usage
from litellm import asearch
import os, asyncio
os.environ["PERPLEXITYAI_API_KEY"] = "pplx-..."
async def search_async():
response = await asearch(
query="machine learning research papers",
search_provider="perplexity",
max_results=10,
search_domain_filter=["arxiv.org", "nature.com"]
)
# Access search results
for result in response.results:
print(f"{result.title}: {result.url}")
print(f"Snippet: {result.snippet}")
asyncio.run(search_async())
Optional Parameters
response = search(
query="AI developments",
search_provider="perplexity",
# Unified parameters (work across all providers)
max_results=10, # Maximum number of results (1-20)
search_domain_filter=["arxiv.org"], # Filter to specific domains
country="US", # Country code filter
max_tokens_per_page=1024 # Max tokens per page
)
LiteLLM AI Gateway Usage
LiteLLM provides a Perplexity API compatible /search endpoint for search calls.
Setup
Add this to your litellm proxy config.yaml
model_list:
- model_name: gpt-5.6-terra
litellm_params:
model: gpt-5.6-terra
api_key: os.environ/OPENAI_API_KEY
search_tools:
- search_tool_name: perplexity-search
litellm_params:
search_provider: perplexity
api_key: os.environ/PERPLEXITYAI_API_KEY
- search_tool_name: tavily-search
litellm_params:
search_provider: tavily
api_key: os.environ/TAVILY_API_KEY
- search_tool_name: parallel-search
litellm_params:
search_provider: parallel_ai
api_key: os.environ/PARALLEL_API_KEY
Start litellm
litellm --config /path/to/config.yaml
# RUNNING on http://0.0.0.0:4000
Test Request
Option 1: Search tool name in URL (Recommended - keeps body Perplexity-compatible)
curl http://0.0.0.0:4000/v1/search/perplexity-search \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "latest AI developments 2024",
"max_results": 5,
"search_domain_filter": ["arxiv.org", "nature.com"],
"country": "US"
}'
Option 2: Search tool name in body
curl http://0.0.0.0:4000/v1/search \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"search_tool_name": "perplexity-search",
"query": "latest AI developments 2024",
"max_results": 5
}'
Load Balancing
Give multiple search tools the same search_tool_name to load balance across them. Each request picks one of the matching tools at random. router_settings.routing_strategy does not apply to search tools, so strategies like least-busy or latency-based-routing have no effect on which provider serves a search request
search_tools:
- search_tool_name: my-search
litellm_params:
search_provider: perplexity
api_key: os.environ/PERPLEXITYAI_API_KEY
- search_tool_name: my-search
litellm_params:
search_provider: tavily
api_key: os.environ/TAVILY_API_KEY
- search_tool_name: my-search
litellm_params:
search_provider: exa_ai
api_key: os.environ/EXA_API_KEY
- search_tool_name: my-search
litellm_params:
search_provider: brave
api_key: os.environ/BRAVE_API_KEY
Test with load balancing:
curl http://0.0.0.0:4000/v1/search/my-search \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "AI developments",
"max_results": 10
}'
Restrict Search Tool Access
Set search_tools under object_permission on a key or team to limit which search tools it can call. The allowlist applies to /search, /v1/search, /search/{search_tool_name}, web search interception, router fallbacks between search tools, and /search_tools/list
curl http://0.0.0.0:4000/team/new \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"team_alias": "research",
"object_permission": {"search_tools": ["tavily-search"]}
}'
By default an empty or missing search_tools list allows every search tool. To make every search tool opt-in, turn on search_tool_deny_by_default:
general_settings:
search_tool_deny_by_default: true
The setting defaults to false. With it on, the requested search tool must be listed in object_permission.search_tools of each identity the request resolves to
| Caller | Grants required |
|---|---|
| Virtual key without a team | The key |
| Virtual key with a team | The key and its team |
Team member without a virtual key (JWT or lite login session) | The team the request resolved to |
| User without a virtual key or team | The user |
A missing permission record, a null list, and an empty list all grant nothing. A user's personal grants only count when the request has no virtual key and no team, so they never widen or narrow a key or team request. If a key names a team that cannot be loaded, the request is denied rather than treated as a key without a team. Denied requests return 403 before any search provider is called, with key_search_tool_access_denied, team_search_tool_access_denied, or user_search_tool_access_denied, and /search_tools/list only returns the tools the caller may call
For a team key, grant the tool on both objects:
curl -X POST 'http://localhost:4000/team/new' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"team_alias": "research", "object_permission": {"search_tools": ["tavily-search"]}}'
curl -X POST 'http://localhost:4000/key/generate' \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"team_id": "<team_id from above>", "object_permission": {"search_tools": ["tavily-search"]}}'
The master key and dashboard login sessions are not restricted. A proxy admin calling with its own virtual key is restricted like any other key. Web search interception with no registered search tool stops falling back to the default provider for every restricted caller, since there is no tool name a grant could list. Existing keys and teams with empty lists lose search access as soon as the setting is on
Setting the flag back to false restores the earlier behavior, where an empty or unset list means unrestricted. A nonempty search_tools list that leaves out the requested tool is still rejected
Request/Response Format
LiteLLM follows the Perplexity Search API specification.
See the official Perplexity Search documentation for complete details.
Example Request
{
"query": "latest AI developments 2024",
"max_results": 10,
"search_domain_filter": ["arxiv.org", "nature.com"],
"country": "US",
"max_tokens_per_page": 1024
}
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string or array | Yes | Search query. Can be a single string or array of strings |
search_provider | string | Yes (SDK) | The search provider to use: "perplexity", "tavily", "parallel_ai", "exa_ai", "brave", "google_pse", "dataforseo", "firecrawl", "searxng", "linkup", "duckduckgo", "searchapi", "serper", or "you_com" or "apiserpent" or "agentcore" or "bing_grounding" |
search_tool_name | string | Yes (Proxy) | Name of the search tool configured in config.yaml |
max_results | integer | No | Maximum number of results to return (1-20). Default: 10 |
search_domain_filter | array | No | List of domains to filter results (max 20 domains) |
max_tokens_per_page | integer | No | Maximum tokens per page to process. Default: 1024 |
country | string | No | Country code filter (e.g., "US", "GB", "DE") |
Query Format Examples:
# Single query
query = "AI developments"
# Multiple queries
query = ["AI developments", "machine learning trends"]
Response Format
The response follows Perplexity's search format with the following structure:
{
"object": "search",
"results": [
{
"title": "Latest Advances in Artificial Intelligence",
"url": "https://arxiv.org/paper/example",
"snippet": "This paper discusses recent developments in AI...",
"date": "2024-01-15"
},
{
"title": "Machine Learning Breakthroughs",
"url": "https://nature.com/articles/ml-breakthrough",
"snippet": "Researchers have achieved new milestones...",
"date": "2024-01-10"
}
]
}
Response Fields
| Field | Type | Description |
|---|---|---|
object | string | Always "search" for search responses |
results | array | List of search results |
results[].title | string | Title of the search result |
results[].url | string | URL of the search result |
results[].snippet | string | Text snippet from the result |
results[].date | string | Optional publication or last updated date |
Supported Providers
| Provider | Environment Variable | search_provider Value |
|---|---|---|
| Perplexity AI | PERPLEXITYAI_API_KEY | perplexity |
| Tavily | TAVILY_API_KEY | tavily |
| Exa AI | EXA_API_KEY | exa_ai |
| Brave Search | BRAVE_API_KEY | brave |
| Parallel AI | PARALLEL_AI_API_KEY | parallel_ai |
| Google PSE | GOOGLE_PSE_API_KEY, GOOGLE_PSE_ENGINE_ID | google_pse |
| DataForSEO | DATAFORSEO_LOGIN, DATAFORSEO_PASSWORD | dataforseo |
| Firecrawl | FIRECRAWL_API_KEY | firecrawl |
| SearXNG | SEARXNG_API_BASE (required) | searxng |
| Linkup | LINKUP_API_KEY | linkup |
| Serper | SERPER_API_KEY | serper |
| DuckDuckGo | DUCKDUCKGO_API_BASE | duckduckgo |
| SearchAPI.io | SEARCHAPI_API_KEY | searchapi |
| You.com | YOUCOM_API_KEY (optional — omit for keyless free tier) | you_com |
| APISerpent | APISERPENT_API_KEY | apiserpent |
| Bedrock AgentCore | AGENTCORE_GATEWAY_URL (required), AWS credentials or AGENTCORE_GATEWAY_TOKEN | agentcore |
| Nimble | NIMBLE_API_KEY | nimble |
| Grounding with Bing (Microsoft Foundry) | BING_GROUNDING_PROJECT_ENDPOINT, BING_GROUNDING_MODEL (required), api_key or BING_GROUNDING_TOKEN or azure-identity | bing_grounding |
See the individual provider documentation for detailed setup instructions and provider-specific parameters.