# PR-001: Public Pre-Registration, deployment pipeline v0 (for a Certificate of Conformance to Pre-Registration PR-001) **Posted by:** Aniket Mallick. **Posting date:** 24 September 2026. This was planned for 10 September; the 14-day slip is stated here, not hidden. **Practice:** Toceta is the certification practice. The workshop repository is ARM-ANI: https://github.com/aniketmallick/robotics_research_arm_ani **Frozen at:** git tag `prereg-001`. This file's SHA-256 and the tag's commit are posted beside the file, because a file cannot hold its own hash. **Ratified by:** the architect channel (RULING-2026-09-24c), against this file's SHA-256. The stamp is posted beside the file with the hash. **In one sentence:** I will take one boring robot task through a five-gate pipeline and publish a *Certificate of Conformance to this pre-registration* on 30 September 2026, pass or fail. Every bar, tail, exclusion rule and "not covered" clause is written here first. ## 1. What "certified" means here (and what it does not) "Certified" means every gate's pre-registered checks were run on named artifacts, and the results are published as measured. It is not a safety certification. It is not a performance guarantee outside the measured envelope. It is not conformance to any external standard. A FAIL certificate is published as a FAIL. **Usage rule (ratified 8 September 2026):** in every public document, post or video from this project, the first use of "certified" links to or restates this definition. The word never appears in a headline without its qualifier. The only form is "Certificate of Conformance to Pre-Registration PR-001". ## 2. The task (boring on purpose) **Pick-and-place of a known cube into a fixed zone.** - **Arm:** SO-101 follower arm (STS3215 servos, 12 V). - **Camera:** one fixed third-person RGB camera (Logitech C920, 640×480, centre-square crop to 96×96). It is the only visual policy input. - **Object:** one 40 mm matte cube, placed on one of 8 printed crosses in a 12×12 cm spawn box, at a yaw from the frozen list (§4a). - **Zone:** an 8×8 cm place zone at a fixed position. - **Lighting:** declared lighting (§2b). **Control.** - Target-delta joint control: at most 0.05 rad per arm joint per step, and at most 0.2 rad of the gripper joint (≈ 10 %) per step. - 15 Hz on the real arm (the lerobot env rate). The sim steps 0.065 s (15.38 Hz), the nearest multiple of its 5 ms physics step. - A trial is at most **100 control steps**. **Success.** Every trial starts from the calibrated rest pose. Success requires all three of these at episode end: - the cube is resting inside the zone; - the gripper is open; - the arm is back within ε = 0.15 rad (∞-norm) of its calibrated rest pose. The definition is identical in sim (G0.4). On the real arm it is read as follows: - **When the episode ends:** at step 100, or earlier once the arm has been within ε of rest with the gripper open for 5 consecutive steps. This is read from the encoders (the sim's 5-step hold). - **"Inside the zone":** the cube's centre is inside the printed 8×8 cm zone outline, as in the sim. - **How "inside the zone" is read:** on a capture taken after a scripted clear-view move (logged as a scoring step outside the trial window), and by the operator's physical read. A trial is a success only if both say so. Every disagreement is reported. - **Occlusion:** at rest the gripper hides the zone's near corner from the camera. This is a declared property of the rest pose. No result is ever scored from a frame in which the gripper hides the zone. **Policy of record.** The certified policy is C1 if it exists, else C0, else NULL-001. C0/C1 are HIL-SERL (SAC) checkpoints trained on the real arm with human takeover through the leader arm, G1-provenanced, selected per §6. NULL-001 is the pre-registered scripted baseline (§4b), which the ruling of 22 September made the certified object of this pre-registration; it remains the certified object unless the readiness gate below passes. **Readiness gate** (pre-registered; decided by 26 Sep 12:00 IST, logged either way). All four must hold: - (i) G3 live with the fault-injection table published, every monitor seen to fire and not fire; - (ii) the real-arm env wrapper running with leader-arm takeover, a 5-minute teleop dry run through it logged under monitors; - (iii) a reward classifier trained on ≥ 200 labelled frames with held-out accuracy ≥ 90 %, hash recorded; - (iv) the actor/learner loop exchanging at least one checkpoint. Any one missing → the policy of record is NULL-001, HIL-SERL moves to PR-002, and the certificate's first not-covered line states it. The gate's outcome cannot be revisited after 12:00. **Training budget,** disclosed on the certificate whatever the result: real-arm interaction time, online episodes, fraction of steps under human control, longest intervention, optimisation steps, checkpoints screened. **Simulation.** A MuJoCo model of the same rig is built and audited under G0. It is the audit surface and the replay venue for the gap number. It is **not** the source of any policy of record. **Why sim-first changed:** a pre-registered Week-1 sim spike missed its bar (1/100 against ≥ 50 %), and the pre-registered fallback fired. The spike ran 60 000 single-environment steps. Published precedents used 25–40 M steps or 1024 parallel environments, so the spike was about 420–670× under-resourced. It measured our compute plan, not the feasibility of sim-first. Properly resourced sim-first is recorded as a v1 question. **Why this task:** there is direct real-world evidence for this class of task: - **Squint** (2026): SO-101, zero-shot, wrist camera only. Place: 10/10 with a cube and 10/10 with a can (20/20). A. Almuzairee and H. I. Christensen, "Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics", IEEE RA-L 2026, https://arxiv.org/abs/2602.21203 (code: https://github.com/aalmuzairee/squint). - **lerobot-sim2real** (2025): SO-100, zero-shot cube pick, from a fixed third-person RGB camera. S. Tao, https://github.com/StoneT2000/lerobot-sim2real (tutorial: `docs/zero_shot_rgb_sim2real.md` in that repository). HIL-SERL on a real arm is Hugging Face's documented workflow (https://huggingface.co/docs/lerobot/en/hilserl). The novelty budget goes entirely to the certification. ### 2a. Declared perception numbers (pre-registered, measured from a real capture) A 40 mm cube in a badly framed image is a rumour, not a task. The measurements come from the A2b of 23 Sep 15:06 on the real layout v3 print, focus locked (anchor record). | Declared quantity | Bar | Measured | |---|---|---| | Cube pixel extent in the policy image | ≥ 10 px | 19.4 px | | Spawn-box (120 mm) pixel span, **both** axes | ≥ 30 px per axis | 43.8 / 44.0 px | | Spawn-box share of frame area | ≥ 20 % | 20.9 % | | Scale | ≤ 4.0 mm/px | 2.74 mm/px | **Derivation of the 4.0 mm/px bar:** 1. The jaw opening is 76 mm tip to tip at 45° of the gripper joint (the vendored MuJoCo model). Minus the 40 mm cube, that leaves 18 mm of geometric half-clearance. 2. That gives a ±8.1 mm lateral centring tolerance (45 % of the geometric figure, discounted for finger convergence and proprioceptive noise). 3. One pixel of localisation error may cost at most half of that, so the bar is ≤ 4.0 mm per pixel. One gripper model, reconciled: on the calipered scale (A8: 91.3 mm at 63.5 % opening, 1.44 mm per %), 76 mm is ≈ 53 % opening, and the declared grasp opening is 63.75 % (≈ 91 mm), so the derivation's 76 mm is the conservative case. **Outer-column disclosure (C3):** - At crosses 4 and 8 the cube's footprint stays inside the policy crop, with 6.6 / 7.1 mm to spare at ±5 mm placement. - Its top face projects up to **10.1 mm outside** the crop (measured). The layout file's −12.6 mm was the pre-print estimate. - This is a property of the camera geometry, not of placement. - Every G2 trial at those cells carries a `top_face_clipped` flag, so a per-cell effect is visible to any reader. The G2 bar is not adjusted for it. ### 2b. The rig as registered: every target is measured in this frame **Layout v3.** - Spawn box 120 mm, separation 40 mm, zone 80 mm; the workspace is 243 mm end to end. - 8 crosses at u = 15 / 45 / 75 / 105 mm from the zone side and v = 30 / 90 mm from the near edge. - SHA-256 `e6e60f79d1c68753e16265b355bca546f010efd1f84277cec8105ee4df9f45c5`. - Layout v2 was superseded before use, and v1 is dead. **Placement.** - **Tolerance:** ±5 mm, the operator's commitment. - **Per-cross check:** A2b checks each cross ± 20 mm ± 5 mm inside the policy crop on the real capture (smallest margin 6.6 mm, at cross 4). - **Crop margins:** about 16.6 mm on the spawn side and 3.9 mm on the zone side. Zone success is scored against the printed outline, not the crop, so the zone-side margin is a note, not a check. - **Sheets:** fixed to the table. A sheet move voids the block, just like a camera move. **Registration of record: r2.** `registration_2026-09-24_r2.json`, SHA-256 `9fb34cea93f5dddd379fad43c9f0a371b9a1575a7dc1556f0fe8113158605c88`. - It supersedes r1 (`d20bc0842df85bc7e2a7aafffdda86a4a03115ff630c1d09f59bca19bcf8a7cf`). r1 stands as written. - **Why r1 was superseded:** r1's A9 (24 Sep 13:18) passed accuracy (rms 3.3 / max 4.9 mm) but **failed its stability bar**. The arm→layout yaw read 4.90° against 8.79° the day before (−3.89°, bar ±1°). The camera and the sheets had held to ≤ 1.1 px, and the servo EEPROM was unchanged. So the arm frame had turned about 4° against the table in 22 hours: **the robot base was not fixed to the table.** That failed bar stays published exactly as recorded. - **The fix:** the base was clamped (two C-clamps on the back of the base plate), recorded as the corrective action. Then one new A9 was taken as a new measurement. - **A9 of record** (24 Sep 14:48–14:50): from the encoder home, torque off, the fixed jaw tip hand-guided to the 8 printed corners. - Against the camera: rms 2.4 mm, max 4.2 mm, scale +0.58 %. That is inside the pre-declared bar (rms ≤ 10 mm, max ≤ 20 mm, scale within 5 %). It is the arm-frame reference. - The bar as first declared also had a ≤ 3° rotation term. It was removed before the run, by ruling, with the prediction on record, because the yaw is a registration constant and not an error. - **Yaw:** 6.16°. It is a registration constant, applied in the transform. Its source is predominantly the A1 pan homing; a component of order 1–2° is base placement below the ruler's resolution (±0.6° per mm). r2's own cause line is read through that dated erratum. - **Stability:** this is the first yaw after the clamp, so the stability bar restarts here. The first pre-G2 re-touch tests ±1° against r2. - **Property of the record:** the registration record is written only on a complete chain, carries a canonical hash, is never overwritten, and names what it supersedes and why. No research number can reach it, because it copies measured fields and erratum ids only. **Camera calibration.** The intrinsics were fitted with the board shape estimated jointly (calibrateCameraRO). Flat-model fits failed the 1 px bar. The physical anchor agrees to +1.1 %. - A2 of record: 23 Sep 15:00, fx 641.5. Held-out reprojection 0.764 px (≤ 1 px). Calibration SHA-256 `f9f4b9c15386541ddabcacf8262da14651b4641869b983a28b3f8f866e327266`. - **The fx hierarchy (C5):** 1. The binding anchor is physical: tape height × A2b's measured scale gives fx_phys = 634.7. 2. The fit must agree with fx_phys within 5 %. 3. The datasheet prior is a declared reference, not a gate. Physical measurement beats datasheet. - **Order:** A2b, then A2, then A9. Focus is locked (autofocus off, hand test). A camera move voids A2b and A2. **Condition 4 (the rigid-target corroboration): attempted, not met.** - One attempt on a taped cardboard backing failed its own 1 px gate (rms 1.40 px). A fit that fails its own gate can neither corroborate nor supersede, so the board-shape A2 stays of record. - **Diagnosis:** correlated. r against the four paper runs = +0.94, +0.94, +0.90, +0.96 (median +0.94, cut ±0.50). Fixed-pattern share 0.945. - **Reading 1** (pre-declared): the backing was not flat. - **Reading 2** (preferred): a pattern that is 94 % fixed in board coordinates across two backings is most likely in the print. The release-object method refines the pattern points, so it absorbs a print error as it absorbs a dome. - For scale only (not as corroboration), the fx values sit within 1.8 % of each other: 641.5 of record, 634.7 physical, 646.2 from the failed attempt. - A screen-shown pattern separates the two readings in one run (PR-002). **Everything else registered:** - Servo EEPROM equals the calibration file on all 6 joints (read and stored). - Rest-pose clearance is 114 mm (≥ 60). - Declared lighting: ceiling LED, blinds closed. - **Lighting floor:** the paper must read brighter than grey 110 in every re-check (measured 152–153). **Stated limits** (the numbers every result is read against): - **A9 residuals:** rms 2.4 mm, max 4.2 mm. - **A2 gates:** - reprojection ≤ 1 px after the shape fit, and held-out ≤ 1 px; - fx within 5 % of fx_phys, and within 15 % of the lens prior; - dome ≤ 5 mm; - fx/fy within 1 %; - principal point not far off-centre; - real tilts ≥ 20° in 3 of 4 directions; - implied camera height within 10 % of the tape; - A2b's fx within 2 % of A2's. ## 3. The gates Each gate has its own pre-registered checklist, ratified 8 September. These are the SHA-256 prefixes of the frozen files. Dated amendments are in §4c. | Gate | File | SHA-256 (first 16) | |---|---|---| | G0 environment & data audit | `phase0/gates/G0_env_data_audit.md` | `bb9478acbee3e4da` | | G1 training provenance | `phase0/gates/G1_training_provenance.md` | `8afd148c9f19cee9` | | G2 frozen evaluation | `phase0/gates/G2_pre_deployment_eval.md` | `0fc17a579d62fdb3` | | G3 runtime monitors | `phase0/gates/G3_runtime_monitors.md` | `256ad2bbdafcc594` | | G4 certification report | `phase0/gates/G4_certification_report_template.md` | `52b73d54c6dd401e` | **Order of execution (Path B′):** 1. G0. 2. **G3 live, including intervention mode.** 3. G1-provenanced real-arm training. 4. Freeze C0 by hash. 5. G2 real evaluation. 6. Sim replay for the gap. 7. G4. Monitors come before motion, and real-arm training *is* motion. Steps 3–4 run only if the readiness gate (§2) passes. Otherwise the NULL-001 session of 25 Sep is the G2 real evaluation (§4b). **G0 items added since 9 September, before every session that uses the arm frame:** - **Framing re-check.** Camera and sheets within 2.5 px of the A2b of record, plus the lighting floor. PASS/FAIL, logged. - **Base guard.** The encoder home, then a touch of two spawn corners against r2's targets: within 1° and 5 mm, else STOP. - The two-corner touch is the guard. The base-ruler reading beforehand is only a coarse pre-check. It is stated so that nobody later reads "ruler passed" as "base unchanged". - A clamp check (hand-tight, not moved) is on the operator's pre-session list. - All three are logged either way. - If the first pre-G2 re-touch fails ±1°, the fix is a through-bolt or a board the base is screwed to, not a third clamp. - **Hard fails** (they were warnings before): - A0 camera mode (fx depends on it). - A0 bus voltage within 10.5–13.5 V. - A4 loop timing at the declared 15 Hz loop: p95 ≤ 75 ms (≥ 13.3 Hz effective) and max ≤ 85 ms. Measured: p95 70.0 ms (≈ 14.3 Hz effective) and max 72.5 ms. This keeps the loop clear of 10 Hz, where control rate was causal in an earlier stage (0/3 grasps). - **G3.** G3.3's session-start checks include the framing re-check and the base guard from this posting. For a learned policy of record, the matched-block step-0 action shift against the §6 screening block (threshold 5°, evaluated on matched blocks only — the scope condition is part of the monitor) is reported per block, not gating, in PR-001; it gates in PR-002. - **Warnings only, listed on the certificate as such:** the A1 withheld sign and the A8 contact checks. - **Hours rule.** No session that involves motor motion begins after 22:00 IST (the runner refuses 22:00–06:00). G2 runs in daylight. ## 4. The bars (frozen; derivations in G2; no bar moves) **G2.1: sim success of the frozen sim checkpoint (n = 400).** - ~~Bar ≥ 310/400~~. Under Path B′ this **no longer gates anything**. The bar existed to stop a *sim-trained* policy reaching the arm. No policy of record here is sim-trained, so the sim run is the replay that produces the gap number, reported without a bar. - This removal was flagged for ratification on 9 September, not dropped silently (§4c). - Tail: the replay's per-cell breakdown is reported. **G2.2: real success of the policy of record (n = 40).** - The bar, tails and CI apply to whichever it is (C1, C0 or NULL-001). 8 cells × 5 trials. - **Bar: ≥ 22/40.** This is a decision threshold between acceptable and unacceptable: - acceptable is ≥ 70 % true success, which passes with P = 0.985; - unacceptable is ≤ 40 % true success (this rig's prior public imitation result), which passes with P = 0.039. - The Wilson 95 % CI is reported alongside. - **Tails.** Any one overturns a passing centre: - **(a)** any cell at 0/5 (false-overturn rate 1.9 % at a true 70 %); - **(b)** 7 or more consecutive failures in session order (0.5 %); - **(c)** any policy-attributable envelope violation counts as a failure and is reported. **G2.3: real-to-sim gap (C0).** - No bar. It is reported with its CI. - A gap above 30 pp flags G0 coverage as insufficient for v1. - Derivation: 30 pp is the designed allowance (85 % − 70 % = 15 pp) plus the 95 % half-width of the gap estimate at n = 400/40 (14.6 pp). - Tail: the max-gap cell. **G2.4: null baseline NULL-001 (§4b).** One session, two roles: - **Fallback case** (the readiness gate fails): NULL-001 is the policy of record, and its session is the certified evaluation under G2.2's bar, tails and CI. - **Non-fallback case:** the same table is reported as G2.4 with its Wilson CI **beside the A9 residuals it drove on** (rms 2.4 / max 4.2 mm), so a reader can see what kind of floor it is. If it succeeds on more than 20 % of its trials (more than 8 of 40), the *task* (not the bar) is flagged for the architect before the certified evaluation. **Statistical honesty clause:** n = 40 is a bench-time constraint. A true 60 % policy passes 79 % of the time, and a true 50 % policy passes 32 % of the time. The certificate reports the point estimate and the CI, not just PASS/FAIL. **G2.6: sensitivity, reported but not used.** - The decision is recomputed on the first 30 trials: the centre bar is 17/30 (the 55 % bar), tails (b) and (c) apply, and tail (a) is not applied because a cell has only 3–4 trials by then. If the decision flips, the report says so and the n = 40 decision stands. - Success is also reported with the "arm within ε of rest" conjunct relaxed. ### 4a. Frozen artifacts, cited by hash For JSON files the hash is canonical: SHA-256 of the JSON without its own `sha256` field, written by Python `json.dumps(sort_keys=True, separators=(",", ":"))` (ASCII-escaped) and encoded as UTF-8. Each JSON file carries its hash in that field. For other files the hash is SHA-256 of the bytes. | Artifact | File | SHA-256 | |---|---|---| | Layout v3 | `phase0/layout/layout_v3.json` | `e6e60f79d1c68753e16265b355bca546f010efd1f84277cec8105ee4df9f45c5` | | Registration r2 (supersedes r1 `d20bc084…`) | `phase0/registration_2026-09-24_r2.json` | `9fb34cea93f5dddd379fad43c9f0a371b9a1575a7dc1556f0fe8113158605c88` | | Null specification | `phase0/null/NULL-001.md` | `3c8d80bb24547a1ae0dcb67741ed7d9a70b848476495fa7f3dfbf2c21608af32` | | Null planner | `anchor/null_plan.py` | `e544e2d1e87eaf9d9323c25309a547580c552eef60aa4aea90acc2b10029ec05` | | Null plan table (from r2) | `phase0/null/NULL-001_plan_9fb34cea.json` | `2be55a5b7e5dd097e01f1847982fccee48a3940336eea844ada0ef25a98ed0a2` | | G2 decision code | `anchor/g2_decision.py` | `84d2207ac04520f05e04f24be106ecb466756a380d9ab3730ad203693b65955c` | | G2 placement list | `phase0/g2/placements_v1.json` | `09999e5914805f6c12953d9aeabc6faa90d7777a80b1be8edbbe5b24e9dfe6bf` | **Placement list.** - 40 evaluation (cell, yaw) pairs in five interleaved rounds of A1, B1, A2, B2, A3, B3, A4, B4. Row A is the far row (crosses 1–4) and row B is the near row (crosses 5–8), so drift shows as a run, not a cell effect. - Yaw = the first 8 bytes of SHA-256(`"20260924:eval:"`), read big-endian, mod 90°. The seed is the posting date, chosen before any trial and never re-drawn. - A 5-pair screening block for checkpoint selection (§6), disjoint from the 40 by construction: B2@62°, B4@65°, A3@49°, A2@35°, B3@16°. **Decision code.** It decides a trial table exactly as G2 states it: centre, tails (a)–(c), exclusions, no re-run by outcome, frozen checkpoint only, frozen placements in order, and the sensitivities. The four synthetic tables G2.8 names are among its 24 fixtures, all passing: - one 0/5 cell with a passing centre → FAIL-by-tail; - 22/40 → PASS; - 21/40 → FAIL; - a non-frozen checkpoint → refused. For NULL-001 the frozen fingerprint the code checks is the plan table's canonical hash (`2be55a5b…`), so the same code decides its table in either role. All 161 of the project's negative and control fixtures run as expected, and the result is recorded in the anchor record. ### 4b. The null baseline, NULL-001 An open-loop scripted pick-and-place. - **What it knows:** only the cell of each trial (which printed cross). It has no camera, no learned parameters, no feedback and no yaw information; the wrist roll stays at its encoder home. - **Same conditions as the policy:** the same arm, rest pose, control loop, per-step caps, 100-step horizon and success reading. - **How it was planned:** once, from r2, into one fixed joint-target table per cross. Replayed blind. - **Steps used** of the 100, by cross 1–8: 88, 91, 95, 99, 81, 83, 88, 93. - **Checked before it was frozen:** - IK residuals ≤ 1 mm. - Joint limits (calibrated range minus 3°). - Both jaw tips clear a 45°-yawed cube at the tolerance edge on the way down. - The held cube stays ≥ 15 mm over the paper. - The path into rest stays ≥ 60 mm over the zone. - The clear-view pose puts no body in frame. - **When it runs:** 25 Sep, after the base guard: 40 trials on the 40 G2 placements. The NULL-001 session runs under full G2 conditions — frozen placement order, live monitors, exclusion rules, block-void rule, both scoring reads, trial rows in the decision-code format — so that in the fallback case the NULL-001 session is the certified evaluation and is not re-run. In the non-fallback case the same table is reported as G2.4 beside the A9 residuals. One session, two roles, stated. ### 4c. Amendments to the frozen gate files (dated 24 September 2026, ratified with this posting) 1. **G2.1:** the sim bar no longer gates anything under Path B′. Flagged on 9 September (AMENDMENT-002 §G). 2. **G2.4:** the null is NULL-001, on the real arm and the 40 G2 placements (§4b). The sim "reach-and-close-at-random" null and its "expected ≤ 5 %" are withdrawn. Its role follows the policy of record (item 8): - in the fallback case, the certified evaluation under G2.2; - otherwise, reported as G2.4 beside the A9 residuals, with the "> 20 % flags the task" check applied before the certified evaluation. 3. **G2 protocol:** the 8 cells are the 8 crosses of layout v3 (A1–A4 far row, B1–B4 near row). The placements are §4a's list. 4. **G2 checkpoint rule:** C1's deadline "by Sep 22 EOD" becomes "a further training block that finishes before the G2 session". 5. **A9 accuracy bar:** the ≤ 3° rotation term was removed before the run, because the yaw is a registration constant (ruling of 23 September, evening #2). 6. **Control rate:** 15 Hz on the real arm and 15.38 Hz in the sim, stated in §2. The caps are per step, so the null's table replays unchanged at the real rate. 7. **G0 / G3 dates:** as in §7. 8. **Policy of record:** the 22-Sep ruling made NULL-001 the certified object; this posting restores a learned policy as first choice with NULL-001 as the pre-registered fallback, behind the readiness gate in §2. Reason: the anchor completed 24 Sep; HIL-SERL is feasible only if the gate passes, and the gate is decided before any training result exists. ## 5. Exclusion rules (fixed now) **Excluded and re-run immediately on the same placement:** - a hardware fault (servo overload, comms loss, camera dropout), evidenced by the monitor log; - an operator placement error, including a placement whose footprint (cross ± 20 mm ± 5 mm) leaves the policy crop, evidenced by the placement photo. It is not a policy failure and not one of the 40. **Never excluded:** any policy behaviour. **Limits:** at most 8 exclusions per session. A 9th makes the session INCOMPLETE (re-scheduled), not FAIL. No trial is re-run because of its outcome. The decision code refuses a table with two counted trials on one placement. **Block void:** if a framing re-check fails mid-block (the camera or a sheet moved), the block is void. Its trials are re-run after re-registration. A re-frame between blocks is a new registration entry. ## 6. Checkpoint selection rule **C0** is the real-arm training checkpoint with the highest score on a real screen of the 5 fixed screening placements (§4a), run every 50 online episodes. Ties break to the later checkpoint. Screen trials are logged as non-certification trials and never enter the evaluation population. C0 is frozen by hash before the certified evaluation. **The number of checkpoints screened is published beside the selected one.** The screen is a *selector, not an estimate*: choosing the maximum biases the screen's own score upward. Because the blocks are disjoint, the certified n = 40 number is unbiased by that selection. **C1** is a later frozen checkpoint, if a further training block finishes before the G2 session. **The policy of record is C1 if it exists, else C0, else NULL-001 (§2).** C0 and C1 are both evaluated with the full protocol, and both results are published. **Training halt rule:** the learner halts within one checkpoint interval (2 000 optimisation steps) of the last online episode. ## 7. Dates (scope flexes; these do not) | Date | Milestone | Public artifact | |---|---|---| | 18–24 Sep | Rig anchor session: calibration, latency, framing, step responses, gripper, kill switch, arm-frame registration (r2) | This pre-registration | | 24 Sep | This pre-registration posted (14 days after the planned 10 Sep; stated, not hidden) | PR-001 | | 25 Sep | G3 live (morning) → fault-injection demo → base guard code and first base-guard run → NULL-001 session in the afternoon, under full G2 conditions (it is also G3's first live shakedown on real motion). Nothing else on the rig. | Monitor sensitivity / false-alarm table | | 26 Sep | Readiness gate by 12:00 → if PASS, reward-classifier data and HIL-SERL training, screened per §6, until 22:00; if FAIL, the day goes to the G2 harness, block-void handling and the G2.7 conformance map. | Readiness-gate outcome | | 27 Sep | Freeze by hash → G2 in daylight (or, fallback, the NULL-001 table of 25 Sep decided by the frozen code, and 27 Sep goes to the sim G0 audit). | — | | 28–29 Sep | The G0 audit of the sim env, completed before the sim replay; the sim replay for the gap; G4 | Real-to-sim gap, per cell | | 30 Sep | G4 certificate signed; repo, env package, data and video public | Certificate of Conformance to PR-001 | ## 8. What is explicitly not covered - **If the readiness gate fails,** the certificate's first not-covered line is, verbatim: "No learned policy was trained in the pre-registration window; this certificate covers the pipeline on the pre-registered null baseline; the learned policy is PR-002." - Generalization to other objects, lighting, placements, cameras, arms or firmware. - Safety for anyone but the operator. - **If the policy of record is learned,** it was trained on the real arm with human interventions, and the result is partly the operator's. The certificate discloses the training budget (§2): real-arm interaction time, online episodes, fraction of steps under human control, longest intervention, optimisation steps, checkpoints screened. - **Control rate:** 15 Hz declared; the measured loop period is 70.0 ms p95 (≈ 14.3 Hz); the sim replays at 15.38 Hz; the ~7 % mismatch is a stated limit of the gap number. - The simulation's role is audit and gap-replay only, so no claim is made that this task transfers from simulation. - Independent verification. The operator trained, evaluated and signed. This is mitigated by this timestamped post, append-only hash-chained trial rows, and full data release for re-analysis. - Sim fidelity beyond the audited parameters. - The null's own limits, stated in its specification: no yaw information, and thin jaw margins in the worst case. ## 9. Instrument defects caught before any result The audit found its own holes before it found anyone else's. Each item has a fix and a fixture. **One night: three tool failures.** A2b measured the operator instead of the outline; A9 lost a correction on re-run; A2 passed a degenerate fit. All three were found by the channel, fixed with a dated erratum and a test, and none reached a result. That is the audit doing its job on its own instruments, eight days before it does it on a policy. **Accepted garbage.** Five anchor steps accepted garbage: - a frozen stream; - a seized joint; - a stuck readback; - a 250° slide angle; - a jaw margin under the floor. **Gates and records:** - A dry-run A10 satisfied the kill-switch gate, so a rehearsal could have unlocked motion. - A refused write lost its block. - The registration's certificate line printed "stable" for a failed stability bar. Fixed before r1 was written. - The null's first draft aimed the fixed jaw tip at the cube centre, so the jaw would have come down on the cube. Every null trial would have failed by construction, and the null would have scored 0. Caught while planning, before anything was frozen or run. A jaw-clearance check now refuses such a plan. **Five evidence-loss defects:** - the A9 sign erratum; - A2 frames of two runs mixed; - A2b photos; - A3 photos; - the A0 EEPROM dump. The fifth is the expensive one. The 09-18 13:06 A0 was the only EEPROM read before the anchor re-homed the servos, and it was overwritten before the step history existed. It could have dated a change to this rig observed before the anchor session; that onset is now undatable for a second, self-inflicted reason. Every re-run now keeps the previous raw files. The record-keeping was tested by the audit before the audit tested a robot. An audit that hides the evidence it lost is the thing it exists to replace. ## 10. Due before G2 (named; they gate G2 on 27 Sep, not this posting) 0. **Readiness gate (§2):** decided by 26 Sep 12:00 IST and logged either way, with its four pieces of evidence (the G3 fault-injection table, the teleop dry-run log, the reward classifier's hash and held-out accuracy, the first exchanged checkpoint). Its outcome cannot be revisited after 12:00. 1. **G0 base guard code:** two-corner touch ≤ 1° / ≤ 5 mm, ruler pre-check, clamp check, logged, seen to fail in both directions. Due before the first NULL-001 session. 2. **G3 monitors live, with the fault-injection demo.** Due before any motion session: the NULL-001 session of 25 Sep is its first live shakedown. 3. **NULL-001 replay runner:** it replays the frozen table only, refuses any other plan or registration hash, and has its own fixtures. 4. **A4 hard fail in code** (the band in §3). 5. **G2 eval harness:** - it writes the trial rows the decision code reads, with `top_face_clipped` at cells 4 and 8; - block-void handling in the harness and in the decision code (§5); - the G2.7 conformance map: every MUST in G2 mapped to a code path. 6. **The G0 audit of the sim env:** before the sim replay. **Done at posting:** the G2 decision code and its synthetic-table fixtures (§4a). ## 11. Change policy This document is frozen at posting. Any change is a dated erratum appended below, ratified by the architect, and the certificate lists all errata. A bar is never moved after a result is seen. --- ### Errata (append-only) _none_