--- name: outlines description: Generates guaranteed-valid structured output from LLMs with Outlines (dottxt.ai), constraining token sampling via finite state machines for JSON schemas, Pydantic models, regex, choice lists, and integer/float types, on Transformers, llama.cpp, vLLM, and limited OpenAI backends. Use when you need output that always parses as valid JSON or matches a regex. Use when extracting typed data into Pydantic models from a local model. Use when classifying text into a fixed set of categories. Use when running batch structured generation on vLLM or Transformers. Use when controlling token sampling at the grammar level. Not for API models needing automatic retries (use Instructor). license: MIT metadata: version: 1.0.0 category: llm-applications maintainer: Kalaris Labs tags: Prompt Engineering, Outlines, Structured Generation, JSON Schema, Pydantic, Local Models, Grammar-Based Generation, vLLM, Transformers, Type Safety dependencies: outlines, transformers, vllm, pydantic --- # Outlines: Structured Text Generation ## When to Use This Skill Use Outlines when you need to: - **Guarantee valid JSON/XML/code** structure during generation - **Use Pydantic models** for type-safe outputs - **Support local models** (Transformers, llama.cpp, vLLM) - **Maximize inference speed** with zero-overhead structured generation - **Generate against JSON schemas** automatically - **Control token sampling** at the grammar level ## Installation ```bash # Base installation pip install outlines # With specific backends pip install outlines transformers # Hugging Face models pip install outlines llama-cpp-python # llama.cpp pip install outlines vllm # vLLM for high-throughput ``` ## Quick Start ### Basic Example: Classification ```python import outlines from typing import Literal # Load model model = outlines.models.transformers("microsoft/Phi-3-mini-4k-instruct") # Generate with type constraint prompt = "Sentiment of 'This product is amazing!': " generator = outlines.generate.choice(model, ["positive", "negative", "neutral"]) sentiment = generator(prompt) print(sentiment) # "positive" (guaranteed one of these) ``` ### With Pydantic Models ```python from pydantic import BaseModel import outlines class User(BaseModel): name: str age: int email: str model = outlines.models.transformers("microsoft/Phi-3-mini-4k-instruct") # Generate structured output prompt = "Extract user: John Doe, 30 years old, john@example.com" generator = outlines.generate.json(model, User) user = generator(prompt) print(user.name) # "John Doe" print(user.age) # 30 print(user.email) # "john@example.com" ``` ## Core Concepts ### 1. Constrained Token Sampling Outlines uses Finite State Machines (FSM) to constrain token generation at the logit level. **How it works:** 1. Convert schema (JSON/Pydantic/regex) to context-free grammar (CFG) 2. Transform CFG into Finite State Machine (FSM) 3. Filter invalid tokens at each step during generation 4. Fast-forward when only one valid token exists **Benefits:** - **Zero overhead**: Filtering happens at token level - **Speed improvement**: Fast-forward through deterministic paths - **Guaranteed validity**: Invalid outputs impossible ```python import outlines # Pydantic model -> JSON schema -> CFG -> FSM class Person(BaseModel): name: str age: int model = outlines.models.transformers("microsoft/Phi-3-mini-4k-instruct") # Behind the scenes: # 1. Person -> JSON schema # 2. JSON schema -> CFG # 3. CFG -> FSM # 4. FSM filters tokens during generation generator = outlines.generate.json(model, Person) result = generator("Generate person: Alice, 25") ``` ### 2. Structured Generators Outlines provides specialized generators for different output types. #### Choice Generator ```python # Multiple choice selection generator = outlines.generate.choice( model, ["positive", "negative", "neutral"] ) sentiment = generator("Review: This is great!") # Result: One of the three choices ``` #### JSON Generator ```python from pydantic import BaseModel class Product(BaseModel): name: str price: float in_stock: bool # Generate valid JSON matching schema generator = outlines.generate.json(model, Product) product = generator("Extract: iPhone 15, $999, available") # Guaranteed valid Product instance print(type(product)) # ``` #### Regex Generator ```python # Generate text matching regex generator = outlines.generate.regex( model, r"[0-9]{3}-[0-9]{3}-[0-9]{4}" # Phone number pattern ) phone = generator("Generate phone number:") # Result: "555-123-4567" (guaranteed to match pattern) ``` #### Integer/Float Generators ```python # Generate specific numeric types int_generator = outlines.generate.integer(model) age = int_generator("Person's age:") # Guaranteed integer float_generator = outlines.generate.float(model) price = float_generator("Product price:") # Guaranteed float ``` ### 3. Model Backends Outlines supports multiple local and API-based backends. #### Transformers (Hugging Face) ```python import outlines # Load from Hugging Face model = outlines.models.transformers( "microsoft/Phi-3-mini-4k-instruct", device="cuda" # Or "cpu" ) # Use with any generator generator = outlines.generate.json(model, YourModel) ``` #### llama.cpp ```python # Load GGUF model model = outlines.models.llamacpp( "./models/llama-3.1-8b-instruct.Q4_K_M.gguf", n_gpu_layers=35 ) generator = outlines.generate.json(model, YourModel) ``` #### vLLM (High Throughput) ```python # For production deployments model = outlines.models.vllm( "meta-llama/Llama-3.1-8B-Instruct", tensor_parallel_size=2 # Multi-GPU ) generator = outlines.generate.json(model, YourModel) ``` #### OpenAI (Limited Support) ```python # Basic OpenAI support model = outlines.models.openai( "gpt-4o-mini", api_key="your-api-key" ) # Note: Some features limited with API models generator = outlines.generate.json(model, YourModel) ``` ### 4. Pydantic Integration Outlines has first-class Pydantic support with automatic schema translation. #### Basic Models ```python from pydantic import BaseModel, Field class Article(BaseModel): title: str = Field(description="Article title") author: str = Field(description="Author name") word_count: int = Field(description="Number of words", gt=0) tags: list[str] = Field(description="List of tags") model = outlines.models.transformers("microsoft/Phi-3-mini-4k-instruct") generator = outlines.generate.json(model, Article) article = generator("Generate article about AI") print(article.title) print(article.word_count) # Guaranteed > 0 ``` #### Nested Models ```python class Address(BaseModel): street: str city: str country: str class Person(BaseModel): name: str age: int address: Address # Nested model generator = outlines.generate.json(model, Person) person = generator("Generate person in New York") print(person.address.city) # "New York" ``` #### Enums and Literals ```python from enum import Enum from typing import Literal class Status(str, Enum): PENDING = "pending" APPROVED = "approved" REJECTED = "rejected" class Application(BaseModel): applicant: str status: Status # Must be one of enum values priority: Literal["low", "medium", "high"] # Must be one of literals generator = outlines.generate.json(model, Application) app = generator("Generate application") print(app.status) # Status.PENDING (or APPROVED/REJECTED) ``` ## Common Patterns Details, code examples and parameter tables: [references/common-patterns.md](references/common-patterns.md). Read it when this step applies. ## Backend Configuration Details, code examples and parameter tables: [references/backend-configuration.md](references/backend-configuration.md). Read it when this step applies. ## Best Practices ### 1. Use Specific Types ```python # ✅ Good: Specific types class Product(BaseModel): name: str price: float # Not str quantity: int # Not str in_stock: bool # Not str # ❌ Bad: Everything as string class Product(BaseModel): name: str price: str # Should be float quantity: str # Should be int ``` ### 2. Add Constraints ```python from pydantic import Field # ✅ Good: With constraints class User(BaseModel): name: str = Field(min_length=1, max_length=100) age: int = Field(ge=0, le=120) email: str = Field(pattern=r"^[\w\.-]+@[\w\.-]+\.\w+$") # ❌ Bad: No constraints class User(BaseModel): name: str age: int email: str ``` ### 3. Use Enums for Categories ```python # ✅ Good: Enum for fixed set class Priority(str, Enum): LOW = "low" MEDIUM = "medium" HIGH = "high" class Task(BaseModel): title: str priority: Priority # ❌ Bad: Free-form string class Task(BaseModel): title: str priority: str # Can be anything ``` ### 4. Provide Context in Prompts ```python # ✅ Good: Clear context prompt = """ Extract product information from the following text. Text: iPhone 15 Pro costs $999 and is currently in stock. Product: """ # ❌ Bad: Minimal context prompt = "iPhone 15 Pro costs $999 and is currently in stock." ``` ### 5. Handle Optional Fields ```python from typing import Optional # ✅ Good: Optional fields for incomplete data class Article(BaseModel): title: str # Required author: Optional[str] = None # Optional date: Optional[str] = None # Optional tags: list[str] = [] # Default empty list # Can succeed even if author/date missing ``` ## Comparison to Alternatives | Feature | Outlines | Instructor | Guidance | LMQL | |---------|----------|------------|----------|------| | Pydantic Support | ✅ Native | ✅ Native | ❌ No | ❌ No | | JSON Schema | ✅ Yes | ✅ Yes | ⚠️ Limited | ✅ Yes | | Regex Constraints | ✅ Yes | ❌ No | ✅ Yes | ✅ Yes | | Local Models | ✅ Full | ⚠️ Limited | ✅ Full | ✅ Full | | API Models | ⚠️ Limited | ✅ Full | ✅ Full | ✅ Full | | Zero Overhead | ✅ Yes | ❌ No | ⚠️ Partial | ✅ Yes | | Automatic Retrying | ❌ No | ✅ Yes | ❌ No | ❌ No | | Learning Curve | Low | Low | Low | High | **When to choose Outlines:** - Using local models (Transformers, llama.cpp, vLLM) - Need maximum inference speed - Want Pydantic model support - Require zero-overhead structured generation - Control token sampling process **When to choose alternatives:** - Instructor: Need API models with automatic retrying - Guidance: Need token healing and complex workflows - LMQL: Prefer declarative query syntax ## Performance Characteristics **Speed:** - **Zero overhead**: Structured generation as fast as unconstrained - **Fast-forward optimization**: Skips deterministic tokens - **1.2-2x faster** than post-generation validation approaches **Memory:** - FSM compiled once per schema (cached) - Minimal runtime overhead - Efficient with vLLM for high throughput **Accuracy:** - **100% valid outputs** (guaranteed by FSM) - No retry loops needed - Deterministic token filtering ## Resources - **Documentation**: https://outlines-dev.github.io/outlines - **GitHub**: https://github.com/outlines-dev/outlines - **Discord**: https://discord.gg/R9DSu34mGd - **Blog**: https://blog.dottxt.co ## See Also - `references/json_generation.md` - Comprehensive JSON and Pydantic patterns - `references/backends.md` - Backend-specific configuration - `references/examples.md` - Production-ready examples ## Agent operating procedure 1. **Check the environment.** Confirm the framework version, model provider, API keys and rate limits. 2. **Pin down the inputs.** Confirm formats, identifiers and parameters from the data or the user. Ask rather than guess any value that changes the result. 3. **Run a small version first.** Test a single call or chain with a known input and inspect raw outputs. 4. **Execute the full task** using the instructions and references above. 5. **Validate the result.** Evaluate on a small labeled set; check structured outputs against their schema; log prompts and responses. 6. **Report.** State what was run (versions, commands, parameters), what was checked, and what is still uncertain. | If this happens | Do this | |---|---| | Outputs do not match the expected schema | Add validation and retries, tighten the schema, or simplify the prompt. | | A function, flag or endpoint in these instructions is missing in the installed version | Check the installed version's own documentation (`help()`, `--help`, official docs), adapt, and tell the user. Never invent an API. | | A required input, identifier or parameter is ambiguous | Ask the user, or state the assumption explicitly before running. | **Integrity rules** - Never fabricate results, parameters, identifiers, citations or statistics. If something cannot be run or verified, say so plainly. - Never send private or sensitive data to external APIs without the user's consent. - Treat version-specific details here as possibly outdated: confirm them against the official documentation for the installed version. - Ask before actions that cost money, consume shared GPUs or cloud quota, touch personal or patient data, or cannot be undone.