openapi: 3.2.0 info: title: LiteLLM Rag API description: 'Proxy Server to call 100+ LLMs in the OpenAI format. **Customize Swagger Docs** 👉 ```LiteLLM Admin Panel on /ui```. Create, Edit Keys with SSO. Having issues? Try ```Fallback Login``` 💸 ```LiteLLM Model Cost Map```. 🔎 ```LiteLLM Model Hub```. See available models on the proxy. **Docs**' version: 1.102.1 tags: - name: Rag paths: /rag/ingest: post: tags: - Rag summary: Rag Ingest description: 'RAG Ingest endpoint - all-in-one document ingestion pipeline. Supports form upload (for files) or JSON body (for URLs). ## Form upload (for files): ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -F file="@document.pdf" \ -F ''ingest_options={"vector_store": {"custom_llm_provider": "openai"}}'' ``` ## JSON body (for URLs): ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "file_url": "https://example.com/document.pdf", "ingest_options": {"vector_store": {"custom_llm_provider": "openai"}} }'' ``` ## Bedrock: ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -F file="@document.pdf" \ -F ''ingest_options={"vector_store": {"custom_llm_provider": "bedrock"}}'' ```' operationId: rag_ingest_rag_ingest_post responses: '200': description: Successful Response content: application/json: schema: {} security: - APIKeyHeader: [] /v1/rag/ingest: post: tags: - Rag summary: Rag Ingest description: 'RAG Ingest endpoint - all-in-one document ingestion pipeline. Supports form upload (for files) or JSON body (for URLs). ## Form upload (for files): ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -F file="@document.pdf" \ -F ''ingest_options={"vector_store": {"custom_llm_provider": "openai"}}'' ``` ## JSON body (for URLs): ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "file_url": "https://example.com/document.pdf", "ingest_options": {"vector_store": {"custom_llm_provider": "openai"}} }'' ``` ## Bedrock: ```bash curl -X POST "http://localhost:4000/v1/rag/ingest" \ -H "Authorization: Bearer sk-1234" \ -F file="@document.pdf" \ -F ''ingest_options={"vector_store": {"custom_llm_provider": "bedrock"}}'' ```' operationId: rag_ingest_v1_rag_ingest_post responses: '200': description: Successful Response content: application/json: schema: {} security: - APIKeyHeader: [] /rag/query: post: tags: - Rag summary: Rag Query description: 'RAG Query endpoint - search vector store, optionally rerank, and generate LLM response. This endpoint: 1. Extracts the query from the last user message 2. Searches the vector store for relevant context 3. Optionally reranks the results 4. Generates an LLM response with the retrieved context ## Example Request: ```bash curl -X POST "http://localhost:4000/v1/rag/query" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "What is LiteLLM?"}], "retrieval_config": { "vector_store_id": "vs_abc123", "custom_llm_provider": "openai", "top_k": 5 } }'' ``` ## With Reranking: ```bash curl -X POST "http://localhost:4000/v1/rag/query" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "What is LiteLLM?"}], "retrieval_config": { "vector_store_id": "vs_abc123", "custom_llm_provider": "openai", "top_k": 10 }, "rerank": { "enabled": true, "model": "cohere/rerank-english-v3.0", "top_n": 3 } }'' ```' operationId: rag_query_rag_query_post responses: '200': description: Successful Response content: application/json: schema: {} security: - APIKeyHeader: [] /v1/rag/query: post: tags: - Rag summary: Rag Query description: 'RAG Query endpoint - search vector store, optionally rerank, and generate LLM response. This endpoint: 1. Extracts the query from the last user message 2. Searches the vector store for relevant context 3. Optionally reranks the results 4. Generates an LLM response with the retrieved context ## Example Request: ```bash curl -X POST "http://localhost:4000/v1/rag/query" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "What is LiteLLM?"}], "retrieval_config": { "vector_store_id": "vs_abc123", "custom_llm_provider": "openai", "top_k": 5 } }'' ``` ## With Reranking: ```bash curl -X POST "http://localhost:4000/v1/rag/query" \ -H "Authorization: Bearer sk-1234" \ -H "Content-Type: application/json" \ -d ''{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "What is LiteLLM?"}], "retrieval_config": { "vector_store_id": "vs_abc123", "custom_llm_provider": "openai", "top_k": 10 }, "rerank": { "enabled": true, "model": "cohere/rerank-english-v3.0", "top_n": 3 } }'' ```' operationId: rag_query_v1_rag_query_post responses: '200': description: Successful Response content: application/json: schema: {} security: - APIKeyHeader: [] components: securitySchemes: APIKeyHeader: type: apiKey description: Bearer token in: header name: x-litellm-api-key