--- name: karpathy-guidelines description: Behavioral guidelines to reduce common LLM coding mistakes. Use when writing, reviewing, or refactoring code to avoid overcomplication, make surgical changes, surface assumptions, and define verifiable success criteria. --- # Karpathy Guidelines Behavioral guidelines to reduce common LLM coding mistakes, derived from [Andrej Karpathy's observations](https://x.com/karpathy/status/2015883857489522876) on LLM coding pitfalls. Use these guidelines to reduce avoidable assumptions and scope drift while still completing authorized work. Investigate evidence and make routine, reversible technical choices within the request; reserve questions for consequential unknowns. ## 1. Think Before Coding **Use evidence. Surface material uncertainty. Make authorized choices.** Before implementing: - Inspect relevant code, tests, configuration, documentation, and the user's stated constraints before deciding what is unclear. - Separate observed facts from assumptions. Make ordinary implementation choices that fit the evidence and project conventions, and summarize consequential choices when useful. - Ask only when an unresolved ambiguity materially changes the outcome, scope, safety, data handling, authorization, or ability to proceed. Do not stop for routine details that can be decided from the available evidence. - Present alternatives when they represent real tradeoffs the user needs to decide; otherwise choose the simplest supported option. - Push back when evidence shows the requested approach will not meet the stated goal or conflicts with a governing requirement. ## 2. Simplicity First **Minimum code that solves the problem. Nothing speculative.** - No features beyond what was asked. - No abstractions for single-use code. - No "flexibility" or "configurability" that wasn't requested. - No error handling for impossible scenarios. - Prefer the smallest change that satisfies the full request and its relevant constraints. Ask yourself: "Would a senior engineer say this is overcomplicated?" If yes, simplify. ## 3. Surgical Changes **Touch only what you must. Clean up only your own mess.** When editing existing code: - Don't "improve" adjacent code, comments, or formatting. - Don't refactor things that aren't broken. - Match existing style, even if you'd do it differently. - If you notice unrelated dead code, mention it - don't delete it. When your changes create orphans: - Remove imports/variables/functions that YOUR changes made unused. - Don't remove pre-existing dead code unless asked. The test: Every changed line should trace directly to the user's request. ## 4. Goal-Driven Execution **Define success criteria. Loop until verified.** Transform tasks into verifiable goals: - "Add validation" → "Enforce the specified invalid-input behavior and verify it with the project-appropriate check" - "Fix the bug" → "Use a focused reproduction or other evidence to confirm the fix" - "Refactor X" → "Verify the behavior that the refactor is meant to preserve" Tests are one verification method, not a universal prerequisite or deliverable. Use the project's required checks and add or run tests when the request, project workflow, or concrete risk calls for them. Test-first is useful for many bug fixes, but do not require deleting valid work or following TDD when the task or project directs another method. For multi-step tasks, state a brief plan: ``` 1. [Step] → verify: [check] 2. [Step] → verify: [check] 3. [Step] → verify: [check] ``` Strong success criteria let you loop independently. Weak criteria ("make it work") require constant clarification. ( # Examples Real-world code examples demonstrating the four principles. Each example shows what LLMs commonly do wrong and how to fix it. --- ## 1. Think Before Coding ### Example 1: Hidden Assumptions **User Request:** "Add a feature to export user data" **❌ What LLMs Do (Wrong Assumptions)** ```python def export_users(format='json'): """Export all users to a file.""" users = User.query.all() if format == 'json': with open('users.json', 'w') as f: json.dump([u.to_dict() for u in users], f) elif format == 'csv': with open('users.csv', 'w') as f: writer = csv.DictWriter(f, fieldnames=['id', 'email', 'name']) writer.writeheader() writer.writerows([u.to_dict() for u in users]) return f"Exported {len(users)} users" ``` **Problems:** - Assumed it should export ALL users (what about pagination? privacy?) - Assumed file location without asking - Assumed which fields to include - Assumed CSV fieldnames without checking actual data structure **✅ What Should Happen (Check Evidence; Ask Only About Material Gaps)** ``` Before implementing, inspect existing export routes, data contracts, authorization rules, and tests. If they do not define the following decisions, ask because they affect privacy and the user-visible result: 1. **Scope**: Export all users or filtered subset? (privacy implications) 2. **Format**: You mentioned "export" - did you mean: - Download file in browser? - Background job with email notification? - API endpoint returning data? 3. **Fields**: Which user fields? (some might be sensitive) 4. **Volume**: How many users typically? (affects approach) Do not invent an endpoint or export every field. If the repository already establishes the format and field policy, follow it. If those consequential decisions remain unspecified, ask only about the missing decisions. ``` ### Example 2: Multiple Interpretations **User Request:** "Make the search faster" **❌ What LLMs Do (Pick Silently)** ```python # Adds caching, database indexes, and async processing # without checking the requested outcome or measuring the current path @lru_cache(maxsize=1000) async def search(query: str) -> List[Result]: # ... 200 lines of optimization code ``` **✅ What Should Happen** ``` First inspect the search path, its current tests and any available measurements. Identify whether the evidence points to response time, throughput, or perceived latency. Do not invent a baseline, target, or effort estimate. If the request and repository evidence support a low-risk optimization, implement and verify it. Ask for a target or priority only when different interpretations would lead to materially different work and no existing requirement resolves the choice. ``` --- ## 2. Simplicity First ### Example 1: Over-abstraction **User Request:** "Add a function to calculate discount" **❌ What LLMs Do (Overengineered)** ```python from abc import ABC, abstractmethod from enum import Enum from typing import Protocol, Union from dataclasses import dataclass class DiscountStrategy(ABC): @abstractmethod def calculate(self, amount: float) -> float: pass class PercentageDiscount(DiscountStrategy): def __init__(self, percentage: float): self.percentage = percentage def calculate(self, amount: float) -> float: return amount * (self.percentage / 100) class FixedDiscount(DiscountStrategy): def __init__(self, fixed_amount: float): self.fixed_amount = fixed_amount def calculate(self, amount: float) -> float: return min(self.fixed_amount, amount) @dataclass class DiscountConfig: strategy: DiscountStrategy min_purchase: float = 0.0 max_discount: float = float('inf') class DiscountCalculator: def __init__(self, config: DiscountConfig): self.config = config def apply_discount(self, amount: float) -> float: if amount < self.config.min_purchase: return 0 discount = self.config.strategy.calculate(amount) return min(discount, self.config.max_discount) # Usage requires 30+ lines of setup for a simple calculation ``` **✅ What Should Happen (Simple)** ```python def calculate_discount(amount: float, percent: float) -> float: """Calculate discount amount. percent should be 0-100.""" return amount * (percent / 100) # Usage discount = calculate_discount(100.0, 10.0) # $10 off ``` **When to add complexity:** Only when you actually need multiple discount types. If that requirement comes later, refactor then. ### Example 2: Speculative Features **User Request:** "Save user preferences to database" **❌ What LLMs Do (Unnecessary Features)** ```python class PreferenceManager: def __init__(self, db, cache=None, validator=None): self.db = db self.cache = cache or InMemoryCache() self.validator = validator or DefaultValidator() def save(self, user_id: int, prefs: dict, merge: bool = True, validate: bool = True, notify: bool = False) -> bool: """Save preferences with optional merging, validation, and notifications.""" if validate: errors = self.validator.validate(prefs) if errors: raise ValidationError(errors) existing = self.get(user_id) if merge else {} if merge: existing.update(prefs) prefs = existing self.db.execute( "UPDATE users SET preferences = ? WHERE id = ?", (json.dumps(prefs), user_id) ) if self.cache: self.cache.set(f"prefs:{user_id}", prefs) if notify: self.notify_preference_change(user_id, prefs) return True def notify_preference_change(self, user_id: int, prefs: dict): # Another 30 lines for a feature nobody asked for pass ``` **✅ What Should Happen (Just What Was Asked)** ```python def save_preferences(db, user_id: int, preferences: dict): """Save user preferences to database.""" db.execute( "UPDATE users SET preferences = ? WHERE id = ?", (json.dumps(preferences), user_id) ) ``` **Add later if needed:** Caching (when performance matters), validation (when bad data appears), merging (when requirement emerges). --- ## 3. Surgical Changes ### Example 1: Drive-by Refactoring **User Request:** "Fix the bug where empty emails crash the validator" **❌ What LLMs Do (Too Much)** ```diff def validate_user(user_data): - # Check email format - if not user_data.get('email'): + """Validate user data.""" + email = user_data.get('email', '').strip() + + # Validate email + if not email: raise ValueError("Email required") - - # Basic email validation - if '@' not in user_data['email']: + if '@' not in email or '.' not in email.split('@')[1]: raise ValueError("Invalid email") - # Check username - if not user_data.get('username'): + # Validate username + username = user_data.get('username', '').strip() + if not username: raise ValueError("Username required") + if len(username) < 3: + raise ValueError("Username too short") + if not username.isalnum(): + raise ValueError("Username must be alphanumeric") return True ``` **Problems:** - "Improved" email validation beyond the bug fix - Added username validation nobody asked for - Changed comments - Added docstring **✅ What Should Happen (Surgical)** ```diff def validate_user(user_data): # Check email format - if not user_data.get('email'): + email = user_data.get('email', '') + if not email or not email.strip(): raise ValueError("Email required") # Basic email validation - if '@' not in user_data['email']: + if '@' not in email: raise ValueError("Invalid email") # Check username if not user_data.get('username'): raise ValueError("Username required") return True ``` **Only changed:** The specific lines that fix empty email handling. ### Example 2: Style Drift **User Request:** "Add logging to the upload function" **❌ What LLMs Do (Reformat Everything)** ```diff - def upload_file(file_path, destination): + def upload_file(file_path: str, destination: str) -> bool: + """Upload file to destination with logging.""" + logger.info(f"Uploading {file_path} to {destination}") + try: - with open(file_path, 'rb') as f: - data = f.read() - - response = requests.post(destination, files={'file': data}) - - if response.status_code == 200: - return True - else: - return False + with open(file_path, "rb") as f: + data = f.read() + response = requests.post(destination, files={"file": data}) + success = response.status_code == 200 + + if success: + logger.info(f"Upload successful: {file_path}") + else: + logger.error(f"Upload failed: {file_path}, status={response.status_code}") + + return success except Exception as e: - print(f"Error: {e}") + logger.exception(f"Upload error for {file_path}: {e}") return False ``` **Problems:** - Changed quote style ('' to "") - Added type hints nobody asked for - Added docstring - Reformatted whitespace - Changed boolean return logic **✅ What Should Happen (Match Existing Style)** ```diff + import logging + + logger = logging.getLogger(__name__) + def upload_file(file_path, destination): + logger.info(f'Starting upload: {file_path}') try: with open(file_path, 'rb') as f: data = f.read() response = requests.post(destination, files={'file': data}) if response.status_code == 200: + logger.info(f'Upload successful: {file_path}') return True else: + logger.error(f'Upload failed: {file_path}, status={response.status_code}') return False except Exception as e: - print(f"Error: {e}") + logger.exception(f'Upload error: {file_path}') return False ``` **Matched:** Single quotes, no type hints, existing boolean pattern, spacing style. --- ## 4. Goal-Driven Execution ### Example 1: Vague vs. Verifiable **User Request:** "Fix the authentication system" **❌ What LLMs Do (Vague Approach)** ``` I'll fix the authentication system by: 1. Reviewing the code 2. Identifying issues 3. Making improvements 4. Testing the changes [Proceeds to make changes without clear success criteria] ``` **✅ What Should Happen (Verifiable Goals)** ``` For "fix authentication," first inspect the reported issue, auth contract, relevant code, and tests. If those sources identify the defect, use that as the success criterion. If no specific failure or intended behavior is available, ask what outcome needs to change. For example, if the reported issue is "users stay logged in after password change": Plan: 1. Confirm the behavior using an existing test or a focused reproduction; add a regression test if it is useful and in scope. 2. Implement: Invalidate sessions on password change according to the existing auth design. 3. Check relevant edge cases, such as multiple active sessions or concurrent changes, when the design and risk make them applicable. 4. Run the focused verification and other project-required gates that apply; report exactly what passed or remains unverified. If the code and issue report leave a consequential behavior decision unresolved, ask only about that decision before proceeding. ``` ### Example 2: Multi-Step with Verification **User Request:** "Add rate limiting to the API" **❌ What LLMs Do (All at Once)** ```python # Implements full rate limiting with Redis, multiple strategies, # configuration system, and monitoring in one 300-line commit # No clear verification steps ``` **✅ What Should Happen (Use the Existing Policy and Verify Its Boundaries)** ``` Inspect the existing API policy, middleware, storage, and tests. Use the configured limits, scope, and storage behavior if they are defined. Verify requests at the allowed limit and just beyond it, plus the relevant reset or concurrency behavior. Do not choose numeric limits, endpoints, an in-memory or distributed backend, or a migration plan without support from the request or project design. If a missing policy changes the implementation materially, ask about that policy; otherwise implement the authorized scope and its integration checks. ``` ### Example 3: Evidence-First Verification **User Request:** "The sorting breaks when there are duplicate scores" **❌ What LLMs Do (Fix Without Reproducing)** ```python # Immediately changes sort logic without confirming the bug def sort_scores(scores): return sorted(scores, key=lambda x: (-x['score'], x['name'])) ``` **✅ What Should Happen (Check the Stated Ordering Contract)** ```python # If ties are required to sort alphabetically, verify that contract: def test_sort_with_duplicate_scores(): """Test sorting when multiple items have same score.""" scores = [ {'name': 'Bob', 'score': 100}, {'name': 'Alice', 'score': 100}, {'name': 'Charlie', 'score': 90}, ] result = sort_scores(scores) assert result == [ {'name': 'Alice', 'score': 100}, {'name': 'Bob', 'score': 100}, {'name': 'Charlie', 'score': 90}, ] # Compare current behavior with the stated tie-break rule. # If it violates the contract, make the smallest supported correction: def sort_scores(scores): """Sort by score descending, then name ascending for ties.""" return sorted(scores, key=lambda x: (-x['score'], x['name'])) # Verify using the project's relevant test or check. ``` --- ## Anti-Patterns Summary | Principle | Anti-Pattern | Fix | |-----------|-------------|-----| | Think Before Coding | Ignores available contract and invents behavior | Check project evidence; ask only about consequential gaps | | Simplicity First | Strategy pattern for single discount calculation | One function until complexity is actually needed | | Surgical Changes | Reformats quotes, adds type hints while fixing bug | Only change lines that fix the reported issue | | Goal-Driven | "I'll review and improve the code" | State the requested outcome and the relevant verification | ## Key Insight The "overcomplicated" examples aren't obviously wrong—they follow design patterns and best practices. The problem is **timing**: they add complexity before it's needed, which: - Makes code harder to understand - Introduces more bugs - Takes longer to implement - Harder to test The "simple" versions are: - Easier to understand - Faster to implement - Easier to test - Can be refactored later when complexity is actually needed **Good code is code that solves today's problem simply, not tomorrow's problem prematurely.** )