# Experiment status and evidence This operational record is kept separate from the book README. It tracks the cross-chapter experiments that require special local implementations, evidence gates, external services, or hardware. Cloning a pinned source repository, installing its dependencies, or passing a smoke test does not establish that an experiment is complete. Statuses in this file describe evidence retained in the repository. A local checkpoint or run directory that is still untracked is useful work-in-progress, but it does not close the clean-clone audit until a reviewable evidence package is committed. Status meanings: - **Complete**: the manuscript's execution and evidence gates have substantive saved evidence. A complete experiment may still produce a negative result. - **Incomplete**: some implementation, execution, or evidence gates remain unsatisfied. - **Reader exercise**: completion requires a reader-operated campaign and its retained evidence rather than a repository checkout alone. This table is a selective operational ledger, not a second numbered index. Each chapter README remains the authoritative ordered experiment list. The WebRTC phone project is listed as an unnumbered add-on and uses a stable, non-numeric evidence identifier. ## Tracked experiments | Experiment | Track | Current status and evidence | | --- | --- | --- | | 4-1 | External-service acceptance | **Incomplete.** The real MCP catalog covers the public-data, multimodal, and filesystem gates, but Google Calendar and Notion remain blocked by missing authorized credentials. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). | | 4-2 | Multimodal comparison | **Complete.** The canonical run compares native vision, extract-to-text, and tool-on-demand processing over the same chart/PDF questions with retained model receipts, latency/usage, tool traces, and an external judge. See the [canonical evidence](../chapter4/multimodal-agent/validation/latest.json). | | 4-3 | External-service acceptance | **Incomplete only at Calendar/email authorization.** The canonical 20-call MCP campaign passes 13/15 gates: core safety, sandboxing, spreadsheet rendering, webhook, headless browser, real GitHub PR mutation, Xvfb Computer Use, and KVM-backed Android execution. Calendar and real email-provider mutations remain unsatisfied, so `official_complete` stays false. See the [canonical manifest](../chapter4/execution-tools/validation/experiment_4_3/real_mcp_gui_20260802T093657Z/manifest.json). | | 4-4 | Human/external-channel acceptance | **Incomplete only at real notification delivery.** The canonical run retains a live repository-user decision on the same pending MCP request, a separate conservative timeout, six real Kimi K3 receipts, 55 MCP calls, and 61 verified manifest files. Email, Telegram, and Slack remain unconfigured and are not simulated, so `official_complete` stays false. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). | | 4-5 | Active tool discovery | **Complete; accuracy uplift not observed.** The 126-tool campaign completed both control and active-discovery arms with 3/3 task success in each arm. Active discovery reduced exposed schema text and elapsed time, while retained detours and malformed actions remain visible. See the [Chapter 4 ledger](../chapter4/EXPERIMENT_LEDGER.md). | | 5-15 | Local implementation | **Incomplete.** The PEDO core, deterministic PostgreSQL demo, and tests are included in [the companion project](../chapter5/permission-embedded-data-objects/). The optional live Agent-generated-code campaign has not been run as a canonical evidence package. | | 5-16 | Local experiment | **Complete; strict joint-advantage hypothesis not observed.** Both Agent-creation arms passed their acceptance gates. The [formal comparison](../chapter5/agent-creator/runs/exp5-12-kimi-k3-20260730-v1/comparison.json) found equal deterministic quality and greater template-arm efficiency, but not strictly higher quality and efficiency together. | | 6-1 | External mailbox experiment | **Incomplete.** The Unipile credential probe returned 401 before mailbox mutation, so the three required real inbound-mail cases have not run. See the [project evidence](../chapter6/agent-with-event-trigger/validation/experiment_6_1/). | | 6-2 | Interruptible asynchronous Agent | **Complete.** All four manuscript scenarios passed with real subprocesses: non-blocking work, queued instruction integration, interruption and recovery, and progress-aware cancellation. See the [canonical summary](../chapter6/async-agent/validation/experiment_6_2/real_subprocess_20260730T052500Z/summary.json). | | 附加 | Local WebRTC speech project | **Complete.** Direct and ReAct arms each pass 20/20 gates over browser-microphone RTP, local Whisper, a real external LLM, TTS, and downlink RTP. PSTN/E.164 is outside the manuscript's local call-user gate. See the [stable-ID manifest](../chapter6/phone-agent/validation/runs/phone-agent-webrtc-audio-20260731-v1/manifest.json). | | 6-5 | Local omni-speech experiment | **Complete; the two paths tie overall with complementary failures.** Pinned MiniCPM-o 4.5 ran locally on one RTX PRO 6000. Native end-to-end and same-model self-cascade each scored 3/4: self-cascade fixed one semantic perception error, while end-to-end preserved speaking-rate evidence erased by the transcript. The [canonical evidence](../chapter6/end-to-end-speech/validation/runs/exp6-5-minicpmo45-20260801-v1/evidence.json) also retains a real 24kHz speech output and passes all acceptance checks. | | 6-6 | Local controllable-TTS experiment | **Complete; the full subjective ordering was not reproduced.** Fish Audio S1 produced the 24-reference library and A/B/C media, and three position-balanced Voxtral listening passes rated the multi-reference arm highest. C > B > A did not reproduce because A outscored B. See the [acceptance evidence](../chapter6/controllable-tts/validation/acceptance.json). | | 6-7 | External reference implementation | **Complete for the bounded read-only task.** The official source and Dockerfile were pinned, the image was built locally with retained image/base digests, and Anthropic returned the requested `claude-sonnet-4-5-20250929` on 16/16 calls. The Agent executed 15 native `computer` actions, did not interact with Google reCAPTCHA, recovered through visible Open-Meteo JSON, and reported 70.2°F / clear sky with `end_turn`. The [canonical acceptance](../chapter6/claude-computer-use-native/validation/runs/exp6-7-anthropic-native-20260803-v2/acceptance.json) passes every source, receipt, action, screenshot, grounding, safety, hash, and credential gate; the historical 401 and two failed task attempts remain separately retained. | | 6-8 | Provider-portable Computer Use | **Complete on the open-model arm.** OpenRouter returned `qwen/qwen3-vl-32b-instruct` for 16/16 real calls; the Agent recovered from a Google CAPTCHA through weather.com and completed in 16 one-action steps. The [canonical evidence](../chapter6/computer-use-open-model/validation/latest.json) retains 15 screenshots, raw responses, the action trajectory, deterministic answer grounding, hashes, and a clean credential scan. | | 6-9 | External hardware track | **Incomplete.** The XLeRobot source and non-actuating preflight are pinned, but no authorized physical teleoperation or book task has run. See the [experiment record](../chapter6/xlerobot-teleoperation/README.md). | | 6-11 | External API and hardware track | **Incomplete.** The exact Gemini Robotics-ER request failed authentication and no robot navigation occurred. A successful planning response, authorized navigation run, and the remaining evidence gates are still required. See the [experiment record](../chapter6/gemini-xlerobot-navigation/README.md). | | 6-13 | External Sim2Real track | **Incomplete.** No local ManiSkill environment, RGB-only PPO checkpoint, >90% simulation evaluation, or real deployment exists. Stages 1–2 require real-scene hardware inputs, stages 3–4 require a suitable GPU environment, and stage 5 requires authorized SO-100 actuation. See the [experiment record](../chapter6/rgb-sim2real-grasping/README.md). | | 7-1 | External benchmark execution | **Complete for the manuscript's bounded five-task campaign.** The pinned τ²-bench telecom run scored 4/5 (Pass@1 0.80), retained all raw trajectories and costs, and traces the failed task to a phone/line mismatch that skipped the required data refuel. Upstream format and trial-count verification pass; full-domain task coverage is explicitly outside this bounded claim. See the [manifest](../chapter7/tau2-bench-eval/validation/runs/exp7-1-openrouter-gpt41mini-telecom-20260802-v1/manifest.json). | | 7-2 | Human benchmark | **Complete.** The retained case set covers easy, medium, and hard tasks from GAIA, AndroidWorld, SWE-bench Verified, τ²-bench, Terminal-Bench, and OSWorld-Verified, with 18/18 first-run trajectories and official verification outcomes. See the [completed case set](../chapter7/experiment-7-2-human-benchmark/README.md). | | 7-3 | Local implementation | **Complete.** The four-grade memory rubric has 60 cases and 180/180 structured judgments with full scope in the [saved evidence](../chapter7/user-memory-system-evaluation/results/full_7_3_structured_rubric_evidence.json). | | 7-4 | Local experiment | **Complete.** JSON Cards, RAG, and hybrid systems produced 180/180 real trajectories across 60 cases, with complete cost and failure analysis in the [saved campaign](../chapter7/user-memory-system-evaluation/results/full_7_4_60_cases_costed.json). | | 7-5 | Local experiment | **Complete.** The known-memory trajectory-prefix campaign ran 33/33 real OpenRouter cells (11 production bad cases × JSON/Markdown/Python-like encodings), with zero API errors and 6/11 deterministic policy passes for each encoding. See the [manifest](../chapter7/user-memory-policy-eval/results/manifest.json) and [report](../chapter7/user-memory-policy-eval/results/policy_prefix_live.json). | | 7-6 | Local experiment | **Complete.** The neutral TTS campaign retained 8/8 content-hashed OpenAI/Fish cells and direct-audio Voxtral judgments across four challenge categories. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). | | 7-7 | Local experiment | **Complete.** The Arena Elo/Bradley–Terry campaign processed 1,799,991 public records, retained rankings, matrices, snapshots, plots, and an independently passing manifest. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). | | 7-8 | Local experiment | **Complete.** The neutral Coding Harness campaign retained 18/18 cells (two models × three tasks × three trials), zero API errors, full trajectories, summaries, and verified artifact hashes in the [manifest](../chapter7/model-action-threshold/results/exp7-8-action-threshold-20260731-v1/manifest.json). | | 7-9 | Local experiment | **Complete.** The eight-turn Agent cost campaign retains four real token/cache/latency arms and the measured KV-cache/context-compression comparison. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). | | 7-10 | Long-running provider benchmark | **Incomplete.** The runner and analyzer exist, but retained evidence contains only 29 smoke/readiness observations: no standard N=100 cells, rate ramp, Agent-cost phase, or 168-hour availability campaign. See the [Chapter 7 ledger](../chapter7/EXPERIMENT_LEDGER.md). | | 7-11 | Local experiment | **Complete.** The full 4 × 3 × 2 × 60 matrix retained 1,440/1,440 real trajectories with zero errors or unpriced usage, complete retrieval/task metrics and factorial analysis, and an independently passing verifier. See the [canonical matrix](../chapter7/user-memory-system-evaluation/results/full_7_11_60_case_matrix.json). | | 7-12 | Emulator evaluation | **Complete evidence; deployment not approved.** The [canonical evidence](../chapter7/android-world/validation/candidate_h5c_api33_local_qwen_20260804/evidence.json) retains all 580/580 unique episodes (116 tasks × five trials), including evaluator failures, with zero runtime errors: 26 strict T3A successes (4.4828%) and mean evaluator reward 0.133621, comprising 77 full-reward states plus one `0.5` partial reward. The completed official Pixel 6/API-33 setup had all 24/24 required apps and ran local Qwen2.5-7B revision `a09a35458c702b33eeacc393d103063234e8bc28` via vLLM 0.19.0 on an RTX PRO 6000 Blackwell 96 GB. Candidate Qwen differs from the paired-source Doubao model, so the result establishes neither same-model uplift nor noninferiority. | | 7-13 | Simulation evaluation | **Complete; action chunking improves a low-success policy.** Pinned OpenVLA-OFT and RoboTwin2 ran two real single-GPU `val_only` arms of 256 episodes each with three RGB views and 14-D proprio/action evidence. Chunk 1 scored 0/256 and chunk 25 scored 26/256 (13/128 IID and 13/128 OOD), a paired +10.15625 pp result. All 486 failures have evidence-backed timeout classifications; the [manifest](../chapter7/openvla-robotwin2-eval/validation/runs/exp7-13-localgpu-20260803-v1/manifest.json) binds 512 rollout-video hashes and passes the strict retained-package verifier. | | 8-6 | Local speech training experiment | **Complete (bounded campaign).** Orpheus and Sesame LoRAs were each trained for 60 optimizer steps on the local RTX PRO 6000, evaluated on held-out examples, and compared against their base arms with 40 retained WAVs. Full adapter identities, hashes, automatic proxy results, and failures are in the [strict report](../chapter8/speech-sft-experiment/validation/exp8-6-20260804-v1/REPORT.md). | | 8-7 | Local multilingual training experiment | **Incomplete.** The SFT implementation exists, but the repository retains no checkpoint or before/after multilingual benchmark. See the [Chapter 8 ledger](../chapter8/EXPERIMENT_LEDGER.md). | | 8-8 | Local training experiment | **Complete.** The retained campaign contains 160/160 training and 80/80 held-out Kimi K3 teacher receipts, a real CUDA-trained SmolLM2-135M-Instruct LoRA checkpoint, and the paired comparison in [`validation/exp8-8-kimi3-smollm2-20260730/`](../chapter8/prompt-distillation/validation/exp8-8-kimi3-smollm2-20260730/). Held-out results: teacher 100%, baseline 0%, trained 95%; ~197× latency speedup; ~75% input-token reduction. All eight evidence gates pass. | | 8-9 | Local training experiment | **Complete; the distillation uplift hypothesis was not supported.** All 24 Kimi K3 teacher cases now retain completed trajectories: 23 passed the deterministic answer verifier and entered SFT, while `aime-2016-9-I` completed under native low-reasoning control with the wrong answer and was correctly rejected. Real CUDA training produced a Qwen2.5-1.5B-Instruct LoRA checkpoint in [`checkpoints/exp8-9-qwen25-1.5b-kimi-k3-20260801-v1/`](../chapter8/cot-distillation/checkpoints/exp8-9-qwen25-1.5b-kimi-k3-20260801-v1/). The [completed three-arm report](../chapter8/cot-distillation/validation/experiment_8_9_complete_20260803_v2.json) records baseline 1/24, student 2/24, teacher 23/24, and nonsignificant paired improvement (p=1.0). | | 8-11–8-16 | External training reproductions | **Incomplete.** The GeneralPoints, V-IRL, SimpleVLA-RL, RLVP, ReTool, and AWorld sources/entrypoints are mapped to varying degrees, but none has a retained full training-and-evaluation campaign satisfying its manuscript gate. See the [Chapter 8 ledger](../chapter8/EXPERIMENT_LEDGER.md). | | 9-8 | External-repository self-evolution experiment | **Complete for the autonomous, review-driven self-update loop; downstream benefit not evaluated.** Pinned Hermes received all ten English chapters and its own source without any supplied candidate gap. It independently chose to add evidence-backed learning signals to persisted trajectories. Three fresh terminal-review rejections were fed back to the original Hermes proposer session; it corrected production-format parsing, persistence-path coverage, and counting-consistency defects until a fourth fresh reviewer returned `VERDICT: ACCEPT`. The accepted patch passes 6 new plus 38 existing focused tests and clean-clone application, but remains unmerged; the proposed downstream ablation campaign was not run. See the [credential-free manifest](../chapter9/hermes-self-evolution/validation/exp9-8-hermes-gpt56luna-autonomous-20260802-v2/manifest.json). | | 9-9 | Longitudinal continual-evolution evaluation | **Complete.** The static, append-only, and evolving arms ran 3 seeds × 14 ordered tasks (126 real model calls). The retained evidence separates transfer, rule replacement, retention, obsolete-rule citation, and paired statistics; only the evolving arm replaces the obsolete 20 kg rule and retains the current 23 kg rule. See the [canonical evidence](../chapter9/self-evolution-eval/validation/latest.json). | | 10-1 | Local architecture comparison | **Complete (bounded comparison).** The repaired Skill arm enforces `load_skill("triage")` before specialist tools while keeping the fixed schema prefix. The canonical v2 campaign retains 30 paired tasks, 12 boundary trajectories, 289 provider receipts, 31 real Tavily receipts, and 60 position-swapped independent judge receipts. Skill passes 15/30 deterministic gates versus Transfer 2/30; Skill is slower and uses more uncached input in this model/configuration. See the [acceptance manifest](../chapter10/multi-role-transfer/validation/comparison/runs/exp10-1-qwen35flash-20260809-v2/manifest.json) and [report](../chapter10/multi-role-transfer/validation/comparison/runs/exp10-1-qwen35flash-20260809-v2/REPORT.md). | | 10-2 | Local multi-agent comparison | **Complete.** The four-role Manager and single-Agent arms translated all 26 units of the retained illustrated/code-heavy technical-book sample, with 12/12 acceptance gates and complete quality, context, token, latency, and resource comparisons. See the [canonical index](../chapter10/book-translation/validation/latest.json). | | 10-3 | External concurrent-agent reproduction | **Complete for the retained Anthropic-caller configuration.** The pinned TalkAct campaign retains 16/16 episodes with no runtime/provider errors and passes all 17 gates. Duplex and strawman tie at 1.0 task success; duplex improves median voice latency from 12.52 s to 2.32 s (5.40×), but strawman has higher probe correctness and lower mean wall time. The invalid Gemini credential required TalkAct's supported Anthropic Sonnet caller override, so this same-family configuration must not be silently pooled with upstream default-Gemini results. See the [acceptance report](../chapter10/talkact-reproduction/validation/runs/exp10-3-talkact-anthropic-caller-20260803-v2/acceptance.json). | | 10-3 | Local WebRTC orchestration experiment | **Complete.** A real LLM autonomously selected the Phone Agent; Playwright, bidirectional RTP, local TTS/Whisper, validation/re-asking, concurrent ask/fill, and one localhost submission pass all 9 gates. PSTN/E.164 is not required by the manuscript. See the [manifest](../chapter10/autonomous-phone-registration/validation/runs/exp10-3-webrtc-raw-20260731-v4/manifest.json). | | 10-5 | External generative-agents reproduction | **Complete; the custom-event diffusion hypothesis was not supported.** The exact pinned 25-persona society completed three 17,280-step, two-virtual-day arms with 148,856 canonical provider receipts and zero logical errors. The custom climate workshop remained limited to its originator, while disabling reflection produced zero evidence-linked reflection thoughts and reduced all four blind plausibility scores; baseline was preferred for 17/25 personas. All 14 gates pass in the [canonical acceptance report](../chapter10/generative-agents/validation/runs/exp10-5-qwen37flash-20260804-v1/acceptance.json). | | 10-6 | Local voice multi-agent experiment | **Complete.** One retained eight-seat v11 game completed three night/day/vote cycles and passed every gate in the same report: six real LLM-tool → macOS `say` → OpenRouter native-audio ASR round trips with exact action agreement, information isolation, a rule-determined winner, and all four strategy criteria. The report retains 13 unique response IDs, 1,650 audio-input tokens, 27 positive-byte TTS events, action history, usage, audio hashes, and judge-attempt provenance; the independent validator rechecked all six audio/action boundaries. See the [canonical report](../chapter10/voice-werewolf/validation/runs/exp10-6-simulated-user-openrouter-20260803-v11/acceptance_report.json). | ## Detailed ledgers The chapter ledgers are the canonical detailed records for acceptance scope, saved evidence, and audit findings: - [Chapter 1 experiment ledger](../chapter1/EXPERIMENT_LEDGER.md) - [Chapter 2 experiment ledger](../chapter2/EXPERIMENT_LEDGER.md) - [Chapter 3 experiment ledger](../chapter3/EXPERIMENT_LEDGER.md) - [Chapter 4 experiment ledger](../chapter4/EXPERIMENT_LEDGER.md) - [Chapter 5 experiment ledger](../chapter5/EXPERIMENT_LEDGER.md) - [Chapter 7 experiment coverage ledger](../chapter7/EXPERIMENT_LEDGER.md) - [Chapter 8 experiment coverage ledger](../chapter8/EXPERIMENT_LEDGER.md) Update this summary whenever one of the tracked completion gates changes. Git history provides the status change log.