--- name: idea-to-blueprint description: "Turn a raw product, app, bot or feature idea into one evidence-backed build blueprint (researched stack, epics, Given/When/Then criteria, tests) that coding agents build one epic per session." category: planning risk: safe source: self source_type: self source_repo: iniesohidham/idea-to-blueprint date_added: "2026-09-30" author: iniesohidham tags: [prd, specification, product-planning, user-stories, acceptance-criteria, tech-stack, agentic-engineering] tools: [claude, codex] license: "MIT" license_source: "https://github.com/iniesohidham/idea-to-blueprint/blob/main/LICENSE" --- # Idea to Blueprint ## Overview Turn a raw idea into a single markdown document that is complete enough for a coding agent to build the product epic by epic, in fresh sessions, without ever having to guess. The blueprint is not a "PRD" in the loose sense: it is the product's only shared memory between the human, the research, and every future build session. The skill runs intake questions, web research and competitor benchmarking, then writes one Markdown blueprint: architecture drivers and style, a version-pinned stack with evidence tags, a quality gate, personas, flows, UX copy, edge cases, and epics broken into stories with Given/When/Then acceptance criteria, tests and a Definition of Done. The blueprint embeds a session protocol so Claude Code or Codex builds it one epic per fresh session. This is the author's own skill, published at [iniesohidham/idea-to-blueprint](https://github.com/iniesohidham/idea-to-blueprint) under the MIT license (see `LICENSE`). The files in `references/`, `assets/` and `scripts/` are unchanged copies of the upstream files; this `SKILL.md` adds the catalog sections (Overview, When to Use, Examples, Limitations, safety notes) around the upstream instructions. ## When to Use This Skill - Use when someone shares a product, app, bot or feature idea (even 2–3 lines) and wants a PRD, spec, backlog, epics, user stories, roadmap or MVP plan, even without saying "PRD". - Use when the user wants deep web research, market research or competitor benchmarking before anything is built. - Use when the user wants an architecture or tech-stack recommendation with justification and pinned versions, or code-quality and review standards for a new product. - Use when the user asks for "a plan for Claude Code/Codex" that a coding agent can execute one epic at a time. - Prefer it over lighter PRD skills when research depth, rigor or zero hallucination matters. - Do not use it to review an existing PRD, to compare two frameworks in isolation, or to write a single piece of UX copy. ## How It Works Run this skill on the strongest model available at the highest effort/thinking setting the harness offers. If you cannot tell whether that is the case, say so once at the start and continue. ### Why this document has to be different A build session that starts from a blank context has exactly two sources of truth: the repo and this blueprint. Whatever is missing, vague, or wrong in the blueprint becomes a guess in code. So the whole skill is organized around three habits: 1. **Evidence or silence.** Every claim about the outside world (a library, a version, a competitor, a regulation, a market figure, a service's availability in a country) carries an evidence tag: `[VERIFIED — , ]`, `[ASSUMED — ]`, or `[UNKNOWN → OQ-nn]` (routed to Open Questions). You never fill a gap with a plausible-sounding guess, because the reader has no way to tell your guess from a fact. 2. **Decide, don't hedge.** "Use X or Y" pushes the decision onto an agent with less context than you have now. Make the call, record the alternatives and why they lost. 3. **Write for zero context.** No "as discussed", no "the usual way". Every ID (persona, flow, screen, copy line, story, criterion) is unique and cross-referenced so a fresh session can resolve everything by search. Everything below serves those three habits. ### Workflow | Phase | What happens | Reference to read | |---|---|---| | 0. Intake | Parse the idea, do a 2–5 search warm-up, ask ONE batch of questions with defaults, wait | `references/intake-questions.md` | | 1. Deep research | Understand the domain, benchmark competitors, derive the architecture drivers and shape, choose the stack for agentic engineering, verify the quality tooling, integrations and locale realities, collect UX benchmarks | `references/research-protocol.md`, `references/architecture-decisions.md`, `references/code-quality-and-review.md` | | 2. Decision brief | Present the interpretation, personas, positioning, architecture drivers and shape, stack, quality gate, epic list and open questions in chat; get a go/correct | (template below) | | 3. Write the blueprint | Write the document section by section, in chunks, into one `.md` file | `references/blueprint-template.md` plus the topic references it points to | | 4. Lint & deliver | Run `scripts/lint_blueprint.py`, fix every error, deliver the file with a short wrap-up | (below) | | 5. Build sessions | The user runs one epic per fresh Claude Code / Codex session using the protocol embedded in the blueprint | `references/session-protocol.md` | Do not skip phases 0 and 2. The user explicitly wants to be asked before a long document is produced, and a two-minute checkpoint prevents a hundred pages built on a wrong interpretation. Do not run the phases out of order: research before the intake answers wastes searches on the wrong problem; writing before the brief is approved wastes the document. Talk to the user in the language they wrote in. Persian in → Persian out for all chat and for the human-facing prose of the document (see Language convention). ### Phase 0 — Intake Read `references/intake-questions.md` before asking anything. The rules that matter most: - **Warm up first.** Run 2–5 quick searches on the idea so the questions are concrete ("I see three products in this space — A, B, C — which is closest to what you mean?") instead of generic. - **One batch, defaults on every question.** Group the questions, put a sensible default next to each, and say that answering "defaults" or skipping any question is fine. Never ask in dribs and drabs. - **Never ask what is already answered or inferable.** Cues from the message decide: language, currency, city names, local services, ".ir" domains, Jalali dates, "کاربر ایرانی" → the users are probably Iranian, so ask to *confirm* it in one line rather than asking "where are your users?". - **Ask only what changes the architecture, the personas, or the scope.** A big product needs about 10–15 questions; a single feature needs 3–6. - Where an elicitation UI tool exists (tappable options), use it for the closed questions; put open questions in prose right after. When the answers arrive, write every unanswered item as an `[ASSUMED — …]` row in your notes; those rows become the Assumptions section later. ### Phase 1 — Deep research and benchmarking Read `references/research-protocol.md` and follow its seven tracks **in order**: - **A. Understand the idea** — domain concepts, vocabulary (this becomes the Glossary), how the problem is solved today, adjacent regulation to flag (never legal advice). - **B. Benchmark** — 5–10 direct and indirect competitors (global, plus local ones when the market is Iran or another specific country), a feature matrix, best-in-class behaviour per key flow, common user complaints (paraphrased from reviews, never quoted), the gap this product fills, a positioning statement. - **C. Architecture drivers and shape** (`references/architecture-decisions.md`) — derive ranked, *measurable* quality-attribute drivers from the personas, the scale ceiling, the failure cost of the core loop and the hosting reality; name the explicit non-drivers; score 2–3 candidate architecture styles against them; decide the shape, the module boundaries and the exit triggers that would justify revisiting it. Verify the framework's own recommended structure and whether boundary rules can be enforced mechanically in the candidate languages. - **D. Stack for agentic engineering** — choose the language, framework, database, testing and tooling using the weighted criteria in the protocol (mainstream + heavily documented, strict static typing, one-command deterministic `check`, fast tests, LTS/stable, few dependencies, hosting fit for the market) **and** the drivers from C. Verify current stable versions, deprecations and licenses, and pin them. - **E. Quality bar and tooling** (`references/code-quality-and-review.md`) — verify the tools that make the gate real for the chosen stack (formatter, strict lint config, typechecker flags, test runner, coverage, mutation testing if a maintained one exists, SAST and dependency audit, secret scanning, duplication detection, accessibility checks, architecture-rule library), the language's official style guide, the CI capability of the chosen host, the current security baseline for the category, and — only if the document will quote them — the current delivery-metric definitions and the current evidence on AI-authored code quality. - **F. Integrations and locale realities** — payments, SMS/OTP, email, maps, push, app stores, hosting/CDN. For Iran, availability of foreign services changes with sanctions and filtering; verify today's status for each one you rely on and record the date. - **G. UX benchmarks** — onboarding, empty states, error handling and accessibility norms from the best products in the space. C before D matters. Choosing the framework first and then writing drivers that happen to justify it is the most common silent failure in a document like this, and it is undetectable in the finished file unless you ask whether any driver could have changed the answer. Log every source (URL, title, access date) as you go; the Sources section is built from this log, not reconstructed afterwards. Expect 20–60 searches for a whole product; post a one-line progress note to the user every few searches ("benchmarking local competitors…", "verifying gateway sandbox…") so a long research phase never looks stalled. If web tools are unavailable, stop and tell the user: a blueprint written from memory alone cannot meet the no-hallucination bar, so either they enable web access or the document is delivered with every external fact marked `[UNVERIFIED]` and a warning on the cover. ### Phase 2 — Decision brief (checkpoint) Before writing the document, post a brief in chat, roughly one screen long, in the user's language: ``` ## Decision brief — **How I read the idea:** 2–3 sentences. What it is, for whom, and the single outcome it delivers. **Users are:** , , . (confirmed / assumed) **Personas:** P1 — one line. P2 … (3–5 max, plus anti-personas if useful) **Benchmark in one paragraph:** who exists, what the best of them do, the gap we take. **Architecture drivers (top 3):** AD-1 … (target), AD-2 …, AD-3 … — and what we are explicitly NOT optimising for. **Shape (locked unless you object):**