--- name: problem-creation description: Guide for creating a single LeetCode problem scaffold - scrape, transform to JSON template, generate files, and verify. Use ONLY when user explicitly requests problem creation via /problem-creation command or provides a specific LeetCode problem number/name. --- # Problem Creation Guide ## Assistant Workflow When user requests a problem by **number** or **name/slug**, the assistant will: 1. **Scrape** problem data using `uv run lcpy scrape` 2. **Transform** data into proper JSON template format 3. **CRITICAL: Include images** - Extract image URLs from scraped data and add to readme_examples with format: `![Example N](image_url)\n\n` before code blocks - Check scraped data for image URLs in the `raw_content` field - Look for patterns: `https://assets.leetcode.com/uploads/...` or `` - Common patterns: `kthtree1.jpg`, `kthtree2.jpg`, `clone_graph.png`, `container.jpg` - Images provide crucial visual context, especially for tree and graph problems - Always verify images are included in `readme_examples` and accessible 4. **Create** JSON file in `src/leetcode_py/cli/resources/leetcode/json/problems/{problem_name}.json` (note the `src/` prefix — the non-`src` path silently fails `bake p-gen` with `Warning: JSON file not found`) 5. **Update tags.json5** - If user specifies tags, manually add problem name to corresponding tag arrays in `src/leetcode_py/cli/resources/leetcode/json/tags.json5` 6. **Generate** problem structure using `bake p-gen` 7. **Update @bakefile.py** - Set `problem = "{problem_name}"` on the `MyBakebook` class to the newly created problem name for easier `bake` command usage 8. **Verify** with `bake lint` - fix template issues in JSON if possible, or manually fix generated files if template limitations 9. **Iterate** if JSON fixes: re-run `bake p-gen -p {problem_name} -f` and `bake lint` until passes to ensure reproducibility **If user does not specify a problem number or name/slug**, run: ```bash uv run python .claude/.dev/next_problem.py ``` This will suggest the next problem to work on from the available problem lists based on completion status. ## Scraping Commands ```bash # Fetch by number uv run lcpy scrape -n 1 # Fetch by slug uv run lcpy scrape -s "two-sum" ``` ### Premium Problems (scraper cannot fetch) Premium problems fail with `Error fetching problem: 'NoneType' object is not iterable`. Fetch the statement from the web instead: - **Primary source**: doocs/leetcode raw markdown mirror — covers ALL LeetCode problems (premium included), no search engine needed. URL is deterministic from the problem number + title: `https://raw.githubusercontent.com/doocs/leetcode/main/solution/0100-0199/0156.Binary%20Tree%20Upside%20Down/README_EN.md` (century folder `0100-0199`, zero-padded number, URL-encoded Title Case) - **Fetch with curl**: `curl -s '' | head -c 4000` — raw.githubusercontent serves static text, no browser needed. Title-guess 404? List the exact folder via GitHub API: `curl -s https://api.github.com/repos/doocs/leetcode/contents/solution/0200-0299 | grep '"name"'` - **Fallback**: `https://leetcode.ca/all/{N}.html` (older wording, usually no constraints) — curl first; Playwright only if a source needs JS. Playwright GOTCHA: `browser_evaluate` HANGS on raw.githubusercontent pages (plain-text doc; eval stalls past 120s) — navigate, then read the auto-saved `.playwright-mcp/page-*.yml` snapshot - Model the JSON on the fetched text, never on memory — statements get revised (e.g. 163 Missing Ranges now returns `list[list[int]]` ranges, not the old `["2", "4->49"]` strings, and is Easy) - Premium problems ship only 2-3 official examples, so most test cases are hand-invented — machine-verify every expectation with a reference implementation before writing the JSON (tree problems: `TreeNode[int].from_list` round-trip) ## JSON Template Format Required fields for `src/leetcode_py/cli/resources/leetcode/json/problems/{problem_name}.json`: **CRITICAL: Use single quotes for Python strings in playground fields to avoid JSON escaping issues with Jupyter notebooks.** **JSON Escaping Rules:** - `playground_test_case`: Use single quotes for string literals (e.g., `s = 'hello'` not `s = "hello"`) - `playground_execution`: Use single quotes for string literals - `playground_assertion`: Use single quotes for string literals - Double quotes in JSON + cookiecutter + Jupyter notebook = triple escaping issues **Test Cases Format:** - `test_cases`: Use structured format with `{"list": ["..."]}` instead of string arrays - Each test case should be a string representation of the tuple/parameters - Example: `{"list": ["('input1', 'input2', expected)", "('input3', 'input4', expected)"]}` **GOTCHA — long string cases hit E501 in generated tests.** Ruff line-length is 100 and the generator emits each parametrize case on its own line (16-space indent + quotes + comma). A case whose full line passes col 100 breaks `p-gen`/lint — BUT only when the overflow contains whitespace: ruff/pycodestyle silently exempt no-whitespace overflow (why a 120-char single-word line elsewhere in the repo passes). Hit with 273 Integer to English Words: cases like `(1234567891, 'One Billion Two Hundred Thirty Four Million ...')` produced 121-char lines. Fix at the JSON level: keep each test_cases entry's string payload under ~80 chars. For long-output problems, pick short-output cases at the same scale instead (`(1000000001, 'One Billion One')` covers the Billion path; drop the 99-char mega cases). Ops-sequence (design) cases that still overflow after trimming the op list: drop ALL spaces after commas (`[[[1,2]],[],[]]` instead of `[[[1, 2]], [], []]`) — still valid Python and buys room. **DO NOT rely on a no-whitespace E501 exemption**: the exemption is inconsistent (batch 11, 3 of 20 agents independently re-tested it: a 102-char no-whitespace parametrize line FAILS `ruff check --select E501`; it holds only when the whole generated line is a single token, and agent reports across runs conflict on the boundary). Worse, `p-gen` runs `ruff format`, which reflows long list literals element-per-line, so compactness does not even survive generation. **Default rule: hard-cap every generated parametrize line at <=100 chars with a write-time assert in the JSON-writing script** (1272/1533 pattern) — if a case cannot fit, use shorter cases at the same coverage instead of compact tricks; 1258 dropped an official example from `test_cases` by trusting the exemption. Compact form is still useful to fit borderline cases, but treat it as unverified until the assert passes. **CAUTION — build compact strings quote-aware**: `repr(x).replace(" ", "")` also strips spaces INSIDE string literals, silently corrupting expectations (`' '` becomes `''`, `'i love you'` becomes `'iloveyou'`; hit twice in one batch, 604 and 642, costing 3 QA cycles). Strip whitespace only OUTSIDE quoted segments (`re.split` on `'...'` spans, strip in non-string parts) or escape literal spaces as `\x20` before stripping. **GOTCHA — every `test_cases` entry must be a COMPLETE Python tuple literal.** The generator parses each entry with `ast.literal_eval`; a bare two-expression string like `'(H)', 'H'` (no enclosing parens) fails to eval and the generator splits it on commas into orphan parametrize values, producing a confusing collection error: `in "parametrize" the number of names (2) must be equal to the number of values (3)`. Parens INSIDE a quoted value are fine — `('Mg(OH)2', 'H2MgO2')` parses correctly; the entry just needs its own enclosing `('...', '...')`. Hit with 726 Number of Atoms: five hand-written paren-formula cases broke collection and cost a full regen cycle. Sanity-check before generating: `python3 -c "import ast; [ast.literal_eval(c) for c in cases]"`. **GOTCHA — the generator does NOT warn when `test_cases` has fewer than 12 entries** (the repo's stated minimum, enforced only at the QA/review level). Count the case list BEFORE running `p-gen` — the JSON-writing script should assert `len(tc) >= 12` alongside its other write-time checks. Hit twice in one batch (723 with 11 cases, 1086 with 6); each needed a post-QA JSON edit + `p-gen -f` regen + solution restore, ~2 extra cycles. When hand-invented cases fall short, top up programmatically: generate random small inputs, run the same reference implementation over them, and assert the outputs into the case list (see 1086 High Five). **IMPORTANT: Create actual JSON files, not JSON5** The template below uses JSON5 format with comments for documentation purposes only. When creating the actual `.json` file, you must: 1. **Remove all comments** (lines starting with `//`) 2. **Use proper JSON syntax** with quoted property names 3. **Save as `.json` file** (not `.json5`) **Template with comments (JSON5 format for reference only):** ````json5 { // ============================================================================ // COMPREHENSIVE LEETCODE TEMPLATE EXAMPLE // ============================================================================ // This example demonstrates ALL template patterns using valid_anagram as base // with comprehensive comments showing variations for different problem types. // // REFERENCE PROBLEMS (see .templates/leetcode/json/ for complete examples): // 1. valid_anagram - Basic: string parameters, boolean return // 2. invert_binary_tree - Tree: TreeNode imports/parameters // 3. merge_two_sorted_lists - LinkedList: ListNode imports/parameters // 4. lru_cache - Design: custom class, multiple methods, operations // 5. implement_trie_prefix_tree - Trie: DictTree inheritance // ============================================================================ // === PROBLEM IDENTIFICATION === problem_name: "valid_anagram", // snake_case: used for directory/file names solution_class_name: "Solution", // "Solution" for basic problems // "LRUCache" for design problems // "Trie(DictTree[str])" for inheritance problem_number: "242", // LeetCode problem number as string problem_title: "Valid Anagram", // Exact title from LeetCode difficulty: "Easy", // Easy, Medium, Hard topics: "Hash Table, String, Sorting", // Comma-separated topics from LeetCode _tags: { list: ["grind-75"] }, // Optional: common problem set tags // Use _tags wrapper for cookiecutter lists // === README CONTENT === // IMPORTANT: Preserve rich HTML content from LeetCode including: // - Code snippets with backticks: `code` // - Bold text: **bold** or bold // - Italic text: *italic* or italic // - Images: ![Example](https://assets.leetcode.com/uploads/...) // - HTML formatting:

,
,