---
title: "/vector_stores/search - Search Vector Store"
url: "/docs/vector_stores/search"
canonical_url: "https://docs.litellm.ai/docs/vector_stores/search"
type: "docs"
last_updated: "2026-10-09"
summary: "Search a vector store for relevant chunks based on a query and file attributes filter. This is useful for retrieval-augmented generation (RAG) use cases."
related:
  - "/docs/vector_stores/create"
  - "/docs/vector_store_files"
---
# /vector_stores/search - Search Vector Store

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


Search a vector store for relevant chunks based on a query and file attributes filter. This is useful for retrieval-augmented generation (RAG) use cases.

## Overview

| Feature | Supported | Notes |
|---------|-----------|-------|
| Cost Tracking | ✅ | Tracked per search operation |
| Logging | ✅ | Works across all integrations |
| End-user Tracking | ✅ | |
| Support LLM Providers | **OpenAI, Azure OpenAI, Bedrock, Vertex RAG Engine, Azure AI, Milvus, Valkey, Gemini** | Full vector stores API support across providers |

For **retrieve, list, update, and delete** over HTTP (including `custom_llm_provider` / `model` routing), see [Create vector store](./create.md#vector-store-management-and-routing-on-the-proxy).

## Usage

### LiteLLM Python SDK

**Basic Usage**

#### Non-streaming example
```python showLineNumbers title="Search Vector Store - Basic"
import litellm

response = await litellm.vector_stores.asearch(
    vector_store_id="vs_abc123",
    query="What is the capital of France?"
)
print(response)
```

#### Synchronous example
```python showLineNumbers title="Search Vector Store - Sync"
import litellm

response = litellm.vector_stores.search(
    vector_store_id="vs_abc123",
    query="What is the capital of France?"
)
print(response)
```

**Advanced Configuration**

#### With filters and ranking options
```python showLineNumbers title="Search Vector Store - Advanced"
import litellm

response = await litellm.vector_stores.asearch(
    vector_store_id="vs_abc123",
    query="What is the capital of France?",
    filters={
        "file_ids": ["file-abc123", "file-def456"]
    },
    max_num_results=5,
    ranking_options={
        "score_threshold": 0.7
    },
    rewrite_query=True
)
print(response)
```

**Multiple Queries**

#### Searching with multiple queries
```python showLineNumbers title="Search Vector Store - Multiple Queries"
import litellm

response = await litellm.vector_stores.asearch(
    vector_store_id="vs_abc123",
    query=[
        "What is the capital of France?",
        "What is the population of Paris?"
    ],
    max_num_results=10
)
print(response)
```

**OpenAI Provider**

#### Using OpenAI provider explicitly
```python showLineNumbers title="Search Vector Store - OpenAI Provider"
import litellm
import os

# Set API key
os.environ["OPENAI_API_KEY"] = "your-openai-api-key"

response = await litellm.vector_stores.asearch(
    vector_store_id="vs_abc123",
    query="What is the capital of France?",
    custom_llm_provider="openai"
)
print(response)
```

**Azure AI Provider**

#### Using Azure AI Search
```python showLineNumbers title="Search Vector Store - Azure AI Provider"
import litellm
import os

# Set credentials
os.environ["AZURE_SEARCH_API_KEY"] = "your-search-api-key"

response = await litellm.vector_stores.asearch(
    vector_store_id="my-vector-index",
    query="What is the capital of France?",
    custom_llm_provider="azure_ai",
    azure_search_service_name="your-search-service",
    litellm_embedding_model="azure/text-embedding-3-large",
    litellm_embedding_config={
        "api_base": "your-embedding-endpoint",
        "api_key": "your-embedding-api-key",
    },
    api_key=os.getenv("AZURE_SEARCH_API_KEY"),
)
print(response)
```

[See full Azure AI vector store documentation](../providers/azure_ai_vector_stores.md)

**Milvus Provider**

#### Using Milvus
```python showLineNumbers title="Search Vector Store - Milvus Provider"
import litellm
import os

# Set credentials
os.environ["MILVUS_API_KEY"] = "your-milvus-api-key"
os.environ["MILVUS_API_BASE"] = "https://your-milvus-instance.milvus.io"

response = await litellm.vector_stores.asearch(
    vector_store_id="my-collection-name",
    query="What is the capital of France?",
    custom_llm_provider="milvus",
    litellm_embedding_model="azure/text-embedding-3-large",
    litellm_embedding_config={
        "api_base": "your-embedding-endpoint",
        "api_key": "your-embedding-api-key",
    },
    milvus_text_field="book_intro",
    api_key=os.getenv("MILVUS_API_KEY"),
)
print(response)
```

[See full Milvus vector store documentation](../providers/milvus_vector_stores.md)

**MongoDB Provider (BETA)**

#### Using MongoDB (BETA)

Search an existing MongoDB Vector Search index on Atlas or a self-managed deployment. Install `litellm[mongodb]`, then set `MONGODB_CONNECTION_STRING` and your embedding provider's credentials. Replace the placeholders with your index, collection fields, and the model used to embed your documents.

```python showLineNumbers title="Search Vector Store - MongoDB Provider (BETA)"
import os

import litellm

response = await litellm.vector_stores.asearch(
    vector_store_id="<index-name>",  # Exact MongoDB Vector Search index name
    query="<question-about-your-documents>",
    custom_llm_provider="mongodb",
    mongodb_connection_string=os.environ["MONGODB_CONNECTION_STRING"],
    mongodb_database="<database-name>",
    mongodb_collection="<collection-name>",
    mongodb_text_field="<text-field>",
    mongodb_embedding_field="<vector-field>",
    litellm_embedding_model="<provider>/<embedding-model>",
    max_num_results=3,
)
print(response)
```

The embedding model must match the one used for the stored vectors. This BETA integration supports search only; index creation, ingestion, filters, ranking options, and query rewriting are not supported.

[MongoDB setup and reference](../providers/mongodb_vector_stores.md) · [Sample-document example](../tutorials/mongodb_vector_search.md)

**Valkey Provider**

#### Using Valkey
```python showLineNumbers title="Search Vector Store - Valkey Provider"
import litellm

response = await litellm.vector_stores.asearch(
    vector_store_id="my-search-index",  # name of the FT index in Valkey
    query="What is the capital of France?",
    custom_llm_provider="valkey",
    valkey_host="my-valkey.example.com",
    valkey_port=6379,
    litellm_embedding_model="openai/text-embedding-3-small",
    max_num_results=3,
)
print(response)
```

[See full Valkey vector store documentation](../providers/valkey_vector_stores.md)

**Gemini Provider**

#### Using Gemini File Search
```python showLineNumbers title="Search Vector Store - Gemini Provider"
import litellm
import os

# Set credentials
os.environ["GEMINI_API_KEY"] = "your-gemini-api-key"

response = await litellm.vector_stores.asearch(
    vector_store_id="fileSearchStores/your-store-id",
    query="What is the capital of France?",
    custom_llm_provider="gemini",
    max_num_results=5
)
print(response)
```

**With Metadata Filter:**
```python showLineNumbers title="Search with Metadata Filter"
response = await litellm.vector_stores.asearch(
    vector_store_id="fileSearchStores/your-store-id",
    query="What is LiteLLM?",
    custom_llm_provider="gemini",
    filters={"author": "John Doe", "category": "documentation"},
    max_num_results=5
)
print(response)
```

[See full Gemini File Search documentation](../providers/gemini_file_search.md)

### LiteLLM Proxy Server

**Setup & Usage**

1. Setup config.yaml

```yaml
model_list:
  - model_name: gpt-5.6-terra
    litellm_params:
      model: openai/gpt-5.6-terra
      api_key: os.environ/OPENAI_API_KEY

general_settings:
  # Vector store settings can be added here if needed
```

2. Start proxy 

```bash
litellm --config /path/to/config.yaml
```

3. Test it with OpenAI SDK!

```python showLineNumbers title="OpenAI SDK via LiteLLM Proxy"
from openai import OpenAI

# Point OpenAI SDK to LiteLLM proxy
client = OpenAI(
    base_url="http://0.0.0.0:4000",
    api_key="sk-<your-litellm-api-key>",  # Your LiteLLM API key
)

search_results = client.beta.vector_stores.search(
    vector_store_id="vs_abc123",
    query="What is the capital of France?",
    max_num_results=5
)
print(search_results)
```

**curl**

```bash showLineNumbers title="Search Vector Store via curl"
curl -L -X POST 'http://0.0.0.0:4000/v1/vector_stores/vs_abc123/search' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
  "query": "What is the capital of France?",
  "filters": {
    "file_ids": ["file-abc123", "file-def456"]
  },
  "max_num_results": 5,
  "ranking_options": {
    "score_threshold": 0.7
  },
  "rewrite_query": true
}'
```

**Vertex AI Search**

Register the data store or search app as a [managed vector store](./managed_vector_stores.md) with provider `vertex_ai/search_api`. Native Discovery Engine search fields go in `extra_body`.

A data store that uses layout-based chunking returns whole-document snippets by default. Ask for chunk results so each hit carries the matching passage:

```bash showLineNumbers title="Search a chunked data store"
curl -L -X POST 'http://0.0.0.0:4000/v1/vector_stores/my-datastore_1234567890/search' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
  "query": "annual pass refund window",
  "max_num_results": 5,
  "extra_body": {
    "contentSearchSpec": {"searchResultMode": "CHUNKS"}
  }
}'
```

Each result's `content[0].text` is the chunk text, `file_id` and `filename` are the source document's URI and title, and `attributes` carry `document_id`, `chunk_id`, `pageSpan`, and the document's `structData` when the store provides them.

An Enterprise-tier search app can return extractive segments or answers instead. They take precedence over snippets in `content[0].text` (segments first, then answers):

```bash showLineNumbers title="Search with extractive content"
curl -L -X POST 'http://0.0.0.0:4000/v1/vector_stores/my-search-app/search' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
  "query": "how does billing work",
  "extra_body": {
    "contentSearchSpec": {
      "extractiveContentSpec": {"maxExtractiveSegmentCount": 1, "maxExtractiveAnswerCount": 1}
    }
  }
}'
```

A structured data store returns each record under `attributes.structData`.

## Setting Up Vector Stores

To search a store that already exists on a provider, register it with LiteLLM first via `config.yaml`, `POST /vector_store/new`, or the Admin UI; see [Managed Vector Stores](./managed_vector_stores.md). For provider-specific configuration, see the [Vector Store Configuration Guide](../completion/knowledgebase.md):

- Provider-specific configuration (Bedrock, OpenAI, Azure, Vertex AI, PG Vector)
- Python SDK and Proxy setup examples  
- Authentication and credential management

## Using Vector Stores with Chat Completions

Pass `vector_store_ids` in chat completion requests to automatically retrieve relevant context. See [Using Vector Stores with Chat Completions](../completion/knowledgebase.md#2-make-a-request-with-vector_store_ids-parameter) for implementation details.

## Related pages

- [/vector_stores - Create Vector Store](https://docs.litellm.ai/docs/vector_stores/create.md)
- [/vector_stores/\{vector_store_id\}/files](https://docs.litellm.ai/docs/vector_store_files.md)
