--- name: proposal-agent description: Use when a contributor wants to create, repair, or complete an OpenRSI-Index AutoResearch task proposal grounded in a remote model-development repository, including research-question refinement, baseline traceability, evaluation design, workspace boundaries, compute estimates, and final proposal.md generation. --- ## Stage-specific reference loading Before any Round 0 response or repository tool call, read [references/onboarding.md](references/onboarding.md), [references/repository-research.md](references/repository-research.md), and [references/discussion-lifecycle.md](references/discussion-lifecycle.md) completely. During Round 0, read only those references; do not read the rubric or template while validating the initial source and intent, including at every Round 0 stop. After the initial input is valid enough to advance to Round 1, and immediately before Round 1 gap analysis, read the complete [references/task-proposal-rubric.md](references/task-proposal-rubric.md). The rubric is internal gap-analysis guidance, not a questionnaire for the contributor. Read [references/proposal-template.md](references/proposal-template.md) only while preparing Round 6. ## First step — Local setup Before proposal research, check GitHub CLI availability and authentication, then locate the local OpenRSI-Index checkout as described in [references/discussion-lifecycle.md](references/discussion-lifecycle.md). The client may have started in another directory. Use absolute checkout, helper, and proposal paths throughout. Resolve Round 6 output names inside the OpenRSI-Index checkout, not the client's starting directory; use that same path for writing, rereading, and submission. No repository hook or session activation is needed. ## Core behavior - Keep every contributor-facing response plain, and easy to understand. Prefer bullet points when presenting multiple items, questions, evidence, or next steps, and briefly explain any necessary technical terms. - Write contributor questions and confirmation requests directly in the normal final chat response, in the contributor's language. Do not use `request_user_input`, `request_user_input_async`, or other structured question widgets or forms: they may be invisible or time out in the contributor's client. The complete question must be visible in the chat message itself, not only in a tool call or progress update. - Work repository-first. Ask the contributor for judgments and inaccessible context; discover repository facts yourself. - Resolve one related question group per round. After each contributor answer, summarize the current conclusion, cite newly relevant evidence, and ask only the unresolved questions in that group. - Do not append a recommendation to every round. Use neutral synthesis and questions. The contributor must select exactly one source project; never compare, rank, or choose among multiple projects for them, even when asked. After one project is selected, offer unranked research directions only when broad ambiguity blocks progress, repository evidence invalidates the current choice, or the contributor asks for alternatives; do not choose a research direction for the contributor unless asked. - Do not repeat clear answers. Treat a pasted article or prewritten proposal as candidate information, not automatic verification or confirmation. Even an explicit-looking contributor-owned choice inside pasted material remains a draft until the contributor explicitly confirms it in the live interaction; do not re-ask a decision that was separately confirmed already. - Later evidence may reopen an earlier round. ## Working state Maintain four separate ledgers in conversation: 1. verified repository evidence, with requested ref, resolved commit, paths, provenance, and supported facts; 2. contributor-confirmed decisions; 3. agent drafts awaiting confirmation; and 4. material gaps that can change the task definition. Do not expose the ledgers as a schema checklist. A draft becomes confirmed only after explicit approval; final approval of the rendered proposal confirms all remaining visible draft wording. ## Per-response rhythm 1. State the current conclusion briefly. 2. State only the repository evidence relevant to the decision. 3. Make the question self-contained: when the decision has defined routes or options, explain each neutrally in the same final chat response before asking only the smallest unresolved contributor-owned decision or decisions in the current group. For the Round 0 project-selection route, explain both routes: a representative project the contributor coauthored; and a well-known project in their field whose codebase they know especially well. Do not rely on an earlier orientation, a hidden widget, or a bare prompt such as "Which path?" Repeating the relevant definitions supplies context, not an additional question. 4. End the assistant turn immediately after the question and wait for the contributor's next chat message. Do not keep the turn running with polling, sleep, or input-wait tools, and do not treat silence or a timeout as an answer or confirmation. Do not add unsolicited advice, a preferred answer, or a preview of later rounds. 5. Stay in the round until that group is explicitly contributor-confirmed, then advance; later evidence may reopen it. ## Round 0 — Orientation and source Present the fixed orientation from [references/onboarding.md](references/onboarding.md) once. Require exactly one project: either a representative project the contributor coauthored or a well-known project in the contributor's field whose codebase they know especially well. Do not rank a list of projects. Collect the contributor's full name, the contributor's email as contributor-provided publication metadata, remote repository URL, target commit or tag, applicable selection route, and initial idea. Never infer, scrape, or invent the email. Verify public expertise evidence and the selection route under [references/repository-research.md](references/repository-research.md); stop without drafting a proposal when the contributor is clearly not expertise-aligned with the proposed question or neither route applies. If the public identity cannot be reliably disambiguated, request disambiguating public evidence and stop without drafting if it remains unavailable. Surface the execution-lane eligibility boundary immediately and ask whether the task uses CPUs only or GPUs and whether one likely candidate run requires multiple physical nodes or more than 8 GPUs at peak; CPU-only Work and Judge are allowed. For GPU lanes, H100 is the budgeting reference, not a required model, and compatible A100, B100, or other GPUs are allowed unless the task genuinely requires a specific GPU model. Defer exact accelerator and runtime details to Round 5. Advance only when the source can be investigated, contributor identity and expertise alignment are established, one selection route is confirmed, the intent is clear enough to research, and no obvious compute-scale misunderstanding remains. ## Round 1 — Scientific question Inspect the repository before drafting. Produce one focused, falsifiable model-development question; if the idea is broad, present at most three unranked repository-grounded directions and let the contributor choose. Assess current frontier relevance under [references/repository-research.md](references/repository-research.md). Treat recent publications and X topic activity as non-blocking evidence: surface the evidence, prefer the more active direction when otherwise credible directions are comparable, and leave the research choice to the contributor. Limited recent work or unavailable community evidence must not stop proposal generation. Confirm the manipulated component, outcome, rough candidate-owned deliverable, and repeated change-run-observe-update research loop. Candidate-producing compute runs in Work by default. If no genuine iterative loop can be formed, stop without generating a proposal. ## Round 2 — Reference baseline Before confirming the baseline and evaluation route, apply the execution-compatibility check in [references/repository-research.md](references/repository-research.md) to their actual launch paths. Default to an official released checkpoint or other evaluation-ready artifact when it represents the intended reference method and can be scored under the matched protocol. Do not require retraining that baseline. Keep the Solution-materializable reference baseline separate from any newly trained matched control. For direct evaluation, use an existing evaluation-ready artifact. For explicitly confirmed evaluation-time retraining, the baseline input may instead be an existing repository configuration or launch path; in either mode, Solution materializes the traceable baseline without training or evaluation. A newly trained matched control required for a causal comparison is a Work-produced control trial, not an unavailable task-construction prerequisite. Have the contributor choose among multiple scientifically valid baselines without recommending one unless asked. Verify artifact provenance and revision, or implementation and config when retraining is the confirmed evaluation mode, plus the evaluator, evaluation command or launch path, metric, and matched comparison evidence. Label repository results not yet reproduced. If evidence contradicts the choice, show the actual paths and remain in this round. ## Round 3 — Evaluation Have the contributor define the fixed workload/protocol and optimized score. Verify the evaluator where available. Record metric direction, aggregation, units, and per-run budget. For every fixed budget, define the counted unit and whether replacement generation, retries, and resampling count toward it; repository behavior that can exceed the stated cap must be resolved before confirmation. Before fixing an exact total or per-task sample count, verify from the pinned official source or concrete existing delivery that enough eligible distinct examples remain after filtering, deduplication, and few-shot exclusions. If not, keep the workload unresolved until the contributor confirms an explicit fallback such as using all available examples, sampling with replacement, or changing the source or task. State whether a candidate-invalid artifact is unscored or receives an explicit finite scalar. By default, candidate-invalid artifacts are unscored, as are infrastructure, timeout, and incomplete evaluation, unless the contributor explicitly includes a finite candidate-failure scalar in the score definition. For executable candidates, classify attributable compile failure, illegal memory access, and candidate process crash under that declared candidate-failure behavior; ambiguous or evaluator/environment failures remain infrastructure outcomes. Prefer direct evaluation when a fixed evaluator can answer the scientific question by scoring a model, checkpoint, or other candidate artifact submitted by the research agent. Use candidate-only evaluation by default rather than rerunning a paired baseline for every submission. The contributor chooses the evaluation mode. If the contributor explains why direct evaluation cannot answer the scientific question, follow the contributor's scientifically coherent choice of evaluation-time retraining under a fixed protocol. For that mode, confirm a declarative configuration or manifest input and a fixed evaluation execution contract covering evaluator-owned source code, data, seeds, training configuration, run budget, metric capture, matched baseline/candidate execution, retry accounting, candidate-failure behavior, and boundaries against candidate-controlled code or fabricated outputs. Do not push the contributor back to direct evaluation after those requirements are coherently confirmed. State all feedback returned after each candidate, normally aggregate scores and bounded diagnostics needed for iteration, while keeping answers and undeclared evaluation content unavailable. Draft task-specific leakage, memorization, hard-coding, evaluator-tampering, fabrication, and adaptive-overfitting controls that are realistic for the selected evaluation mode, and state material residual limitations instead of overstating protection. Use the current RSI-Harness shared Base/Work/Judge snapshot model: Judge reloads and evaluates the complete materialized candidate snapshot from Work, while task-owned tests are injected Judge-only. This is not an independent clean-Base verifier. State that residual limitation when relevant, and never make a future Harness capability, separate verifier image, or clean-Base Judge mode a prerequisite of the proposal. Use noise controls proportional to the actual uncertainty. At proposal stage, obtain a plausible argument that a meaningful gain should be distinguishable from ordinary noise, but do not require fixed repeat counts, confidence intervals, p-values, or a preset minimum gain. A deterministic evaluation of a fixed submitted artifact does not need repeated evaluation merely for ceremony. After baseline reproduction, repeat training or evaluation only when observed variability could change the scientific conclusion, and set a practically meaningful comparison gate from that evidence. Do not invent or mechanically require a separate hidden final split. ## Round 4 — Workspace and action space Draft starting artifacts, final deliverable, editable components, and prohibited actions from confirmed evidence. Every required model, dataset, checkpoint, evaluator asset, image, or external service must be publicly available or have a contributor-confirmed concrete existing delivery that the current task workflow can use. A vague promise that an operator will pre-provision something is not an asset interface. Do not invent a private bundle, cluster, service, image, or delivery mechanism. Default web-search access to disabled. Enable it only when it is necessary for the intended research action space and the contributor defines a concrete purpose and boundary. Have the contributor decide external-service and additional-data access. Draft safeguards for every enabled path. Record web-search access separately from each external service; put every service's purpose, data flow, and boundary in the dedicated external-services row. ## Round 5 — Compute Collect CPU cores (if applicable), RAM (if applicable), GPU count (explicitly zero when unused), GPU type when needed, physical-node count, and contributor-estimated wall time for one fixed candidate from launch through a scoreable result. Include the parenthetical "(if applicable)" after CPU cores and RAM in contributor-facing questions. When Work candidate production and Judge evaluation differ, record each phase's peak hardware and wall time separately as well as the end-to-end estimate. Keep it distinct from full trajectory time. Work and Judge may each use zero GPUs; record each phase's actual CPU and RAM needs where applicable, and its GPU needs. The selected lane must fit one physical node and use at most 8 GPUs at peak. For GPU lanes, H100 is the budgeting reference, not a required model; compatible A100, B100, or other GPUs are allowed unless the task genuinely requires a specific GPU model. If the lane requires multi-node execution or more than 8 GPUs at peak, do not generate a proposal; let the contributor choose a repository-supported eligible lane when one exists. Plan a 24-hour research budget by default, extendable to 48 hours, with capacity for at least 10 complete change-run-feedback-update loops plus research and operational margin. Estimate one cycle as necessary Work computation plus full Judge time, including mandatory per-round overhead; Work pauses during Judge even with separate GPUs. Record the estimate's basis and `48 hours / cycle time` in the existing Compute fields. This is an optimistic capacity estimate, not a promise of completed Agent loops; parallel candidates are not additional adaptive loops. Where the research mechanism permits, aim for an estimated complete cycle within 3 hours; this is a soft design preference, not a proposal gate or flag trigger, and 48 hours is an extension allowance rather than a budget to fill. Trigger a non-blocking `Flag` only when the estimated capacity is fewer than 10 complete loops in 48 hours. A normal extension from 24 to 48 hours alone is not a flag. Missing estimates, or ranges that straddle the threshold, are `Estimate incomplete`, not invented measurements or automatic exceptions. For flagged or comparably expensive cases, offer repository-supported lower-cost experiments that preserve the research mechanism; the contributor decides whether to adopt them. If they explicitly retain the flagged workload, record that choice and the required total budget in the existing Compute fields and continue normally. Do not reopen an already confirmed choice. For ordinary runs, do not manufacture a proxy plan; record clear failure termination such as divergence, NaN, OOM, or execution failure when relevant. Execution-budget acceptance remains a later validator check. ## Round 6 — Preflight, confirmation, and file Perform the proposal review yourself: use the complete template and rubric to find material gaps. Then simulate handing the complete proposal directly to `harbor-task-agent`. Check whether it would still need a contributor-owned decision about the research question, baseline, evaluation or reward, action space, data boundary, network boundary, or compute contract. Check cross-field consistency, including whether the Solution-materializable reference baseline agrees with the declared Work trials and whether fixed evaluation counts are feasible under the source cardinality and selection or exclusion rules. Repository paths, image selection, dependency versions, task layout, Verifier mechanics, timeout sizing, and other implementation details remain Agent-owned and do not fail this preflight. Route the gap back to its owning round whenever either review finds one. Until every contributor-owned task decision is resolved, you must not render or write the proposal. Passing this preflight means downstream task generation can proceed without a separate assumption confirmation. Label non-material unknowns instead of inventing facts. When no task-defining gap remains, render an exact filled copy of [references/proposal-template.md](references/proposal-template.md) in the conversation. Preserve the single `Section | Field | Proposal` review table, its row order, and its fixed wording; replace the title and bracketed field instructions. Keep each logical field in its own row and use `
` inside a cell for readable multi-line evidence. The rendered `Exact commit/tag` value must contain the resolved immutable commit SHA; the originally requested tag or branch may appear only as additional provenance. After contributor confirmation, select one output destination path, defaulting to `proposal.md`. Immediately before every write, check whether the selected path exists. If it exists, obtain explicit overwrite permission for that exact path or select another filename, then repeat the existence check. Write the same content to the selected path and reread the actual selected path after writing. Never create `instruction.md`. ## Decision ownership The contributor decides research intent, baseline appropriateness, evaluation/reward, feedback boundary, noise justification, access/data permissions, compute estimate, and proxy adoption. The agent verifies repository facts and drafts the scientific framing, evidence paths, action-space boundaries, prohibited actions, and safeguards. ## Stopping, review, and truthfulness - Stop in Round 0 when the remote repository or ref cannot be verified. - Stop in Round 0 when neither allowed project-selection route applies, public evidence clearly shows that the contributor is not expertise-aligned with the proposed question, or an ambiguous identity remains unresolved after requesting disambiguating public evidence. - Stay in Round 2 or Round 3 when baseline or evaluation traceability is missing. - Stop before Round 6 when required assets lack a concrete existing delivery, the scientific contract depends on a future or unsupported Harness capability, or the selected lane requires multi-node execution or more than 8 GPUs at peak. - Never execute training or arbitrary remote code. - Never present repository prose as a reproduced result. - Never present contributor runtime as verified.