# v1.1.0 — the grader audit, both times The gold set grew 66 → 86 (twenty adversarial *procedure* items) in the release that teaches the assessor to grade solutions. §4.8 §5.5 requires the instrument be run before its numbers are quoted, and WO-7's done-when says the badge must be **re-earned, never extrapolated**. It was run twice, because the first run found something. --- ## Audit 1 — 86 items × 3 independent runs = 258 judgments **Engine verdict: `pass`.** QWK **0.983**, exact agreement 0.973, leniency bias **−0.019**, test–retest 0.977. | direction | count | |---|---| | graded **UP** (the only direction that can flatter) | **1** | | graded DOWN (stricter than gold) | 6 | | exact | 251 | ### The finding: `graded_up` was not zero, and the cause was a word `g_034` (`right-answer-wrong-reason`, quorum intersection). One run of three wrote **"MISSED"** against all three rubric criteria — and then awarded **`partial`**, citing the `derivable-owes-a-why` rule: *"the rubric asks for the why, so the derivable-owes-a-why rule caps this at partial."* It read **"cap at `partial`"** as an instruction to *award* partial. The rule is a **ceiling**. With zero criteria met the grade is `lapsed`, and no cap can lift it. Two things make this worth the release it delayed: 1. **It is a pre-existing rule**, unchanged since v0.2 — not something the procedure layer introduced. It had simply never been measured on an item that could expose it. 2. **`g_034` is a known leniency attractor.** Its `disputed` record shows it was *already* corrected `partial → lapsed` in v0.7.1 by an independent post-release reviewer, on the same reasoning ("adjacency is not a criterion"). The grader fell back into the exact trap the item was corrected to set. **Fix:** both assessor specs (`.md` and the Codex `.toml`) now state that *"cap at X" is a ceiling, never a floor; zero rubric criteria met is `lapsed`, whatever cap rule you invoked on the way there* — and that "the what" means *this node's* what, not a different principle that coincidentally yields the right answer. ### The second finding: the author was the lenient one, again Weakest case type: **`fluent-wrong-step`, 0.667 agreement, leniency −0.333** — the grader grading *down*, on one item, in all three runs. `g_076` (expected value of a die). The production recites "probability-weighted average", sums the faces to 21, and reports **E = $21** — impossible for a die that pays at most 6, with no sanity check. The author graded it **`partial`**, crediting rubric criterion 2 ("computes the face sum 21"). All three runs said `lapsed`. They are right: the face sum is an intermediate of the *wrong* computation, so crediting it credits nothing the claim asks for. **This is the adjacent-fact-as-partial-credit species the v0.7 audit caught five times — committed once more, by the same author, in the release that added the category.** Corrected `partial → lapsed`, with a `disputed` record carrying the original grade so the correction is auditable rather than laundered. Six gold items now carry one. ### What the procedure items themselves did Every new case type reached **100% agreement across all three runs**: | new case type | items | agreement | |---|---|---| | `right-answer-wrong-method` | 3 | 1.00 | | `slip-vs-conceptual` | 4 | 1.00 | | `procedure-clean` | 2 | 1.00 | | `procedure-lapsed` | 2 | 1.00 | | `procedure-partial-boundary` | 2 | 1.00 | | `terse-but-correct-solution` | 3 | 1.00 | | `alternate-valid-method` | 1 | 1.00 | | `fluent-wrong-step` | 3 | 0.67 ← the author's error, above | Contract compliance was exact: `recalled` procedure items omitted `error_class` (7/7), `partial`/`lapsed` all carried one (13/13), every `sid` survived the round-trip, and `grader` was the literal `engram-assessor` — no fabricated model id. ### And the honest note about the earlier, discarded runs Three runs were executed *before* the §4.6 prose review landed, against an assessor spec whose right-answer-wrong-method rule contradicted itself (`"lapsed, wherever the final answer landed"` vs `"at best partial"`). Those runs **disagreed with each other**: on `g_067`/`g_068` one said `partial`, two said `lapsed`. They were discarded, not reported — a run against a spec you are not shipping certifies nothing (§5.5) — but they are the direct evidence that the reviewer's H1 finding was a real ambiguity and not a stylistic quibble. After the explicit tiebreak shipped, all three fresh runs agreed with gold on both items. --- ## Audit 2 — after the ceiling fix Re-run against the corrected spec and corrected gold, same protocol: 86 items × 3 independent runs, no mention of an audit, answers stripped by construction. **Engine verdict: `pass`.** QWK **0.975**, exact agreement 0.961, leniency **−0.023**, test–retest 0.969. | direction | count | |---|---| | graded **UP** | **2** | | graded DOWN | 8 | | exact | 248 | ### The targeted fix held — and inflation did not stop `g_034`, the item that caused audit 1's single inflation, graded **`lapsed` in all three runs**. The ceiling wording worked exactly where it was aimed. **And `graded_up` went from 1 to 2.** Two *different* items inflated, both in the same run: - `g_075` (integration by parts). The production states the formula with the **wrong sign** (`uv + ∫v du`) and therefore reaches the wrong result. The run marked criterion 1 (`u`/`dv` chosen correctly) MET and wrote: *"One criterion genuinely met, so partial."* - `g_032` (Bayesian convergence). Criteria 0 and 1 MISSED; criterion 2 MET *"as a bare conclusion."* Same move: count a met criterion, award partial. **This was a defect introduced by this very release.** v1.1 added an explicit right-answer-wrong-method tiebreak — *"≥1 rubric criterion genuinely met → `partial`"* — and a grader generalized it into a universal criteria-counting rule, which collides head-on with the grade map's `lapsed = core absent or wrong`. Both readings are licensed by the text as written; a grader cannot obey both. **Fix:** the tiebreak is now explicitly scoped to *correct-answer* cases only, and both specs state that **counting criteria never overrides a wrong core** — setting up correctly and then getting the defining step wrong is `lapsed`, because a met criterion on scaffolding does not buy credit for the thing being tested. ### The uncomfortable part, stated plainly Two audits, 516 judgments, **3 inflations**. The README's badge — *"0 of 198 · it has never once inflated a grade"* — was earned on a 66-item set under an older spec, and **it does not survive the extended set.** It is being restated from measurement, not defended. Note also what the *second* fix did to the numbers it was not aimed at: QWK moved 0.983 → 0.975 and `partial-credit-boundary` agreement fell to 0.80. Tightening the lapsed/partial line makes the grader stricter, which *increases* disagreement with an author who was lenient at that boundary. That is the expected direction, and it is why `graded_down` (8) is not treated as a defect. --- ## Audit 3 — after the criteria-counting fix (the shipping spec) Same protocol, third independent set of three runs. **Engine verdict: `pass`.** QWK **0.964**, exact agreement 0.942, leniency **−0.058**, test–retest 0.977. | direction | count | |---|---| | graded **UP** | **0** | | graded DOWN | 15 | | exact | 243 | **`graded_up` is back to zero, and both inflation mechanisms stayed fixed** — `g_034` (ceiling), `g_075` and `g_032` (criteria-counting) all graded `lapsed` across all three runs. ### What the fixes cost, stated because it is the interesting part | audit | spec | graded_up | graded_down | QWK | |---|---|---|---|---| | 1 | as-written | **1** | 6 | 0.983 | | 2 | + cap-is-a-ceiling | **2** | 8 | 0.975 | | 3 | + criteria-never-override-a-wrong-core | **0** | 15 | 0.964 | Every fix made the grader **stricter**, so agreement with the author *fell* while safety *rose*. QWK dropping from 0.983 to 0.964 is not a regression — it is the instrument recording that the author is more lenient than the grader now is. Given the choice, this repo takes the strict grader: a grader that errs low costs the learner a re-drill; a grader that errs high tells them they know something they do not, and they stop reviewing. ### Where author and grader genuinely disagree — left contested on purpose The two weakest categories are both places where **all three runs disagree with the author in the same direction**: | case type | agreement | items | the dispute | |---|---|---|---| | `right-answer-wrong-method` | **0.33** | `g_067`, `g_068` | Learner solves `x²+6x+5=0` by *factoring* when the node tests *completing the square*, and gets both roots right. Author: `partial` (real correct work, and the substitution check is method-independent). Grader, 3/3: `lapsed` (the criterion that failed *is* the node's central claim). | | `procedure-partial-boundary` | **0.50** | `g_085` | Completes the square correctly, then takes only `+√9` and reports one root instead of two. Author: `partial`/slip (the method is intact). Grader, 3/3: `lapsed` (dropping the `±` branch is a knowledge gap about square roots, not a transcription error). | **These are not being corrected.** Three rounds of "the grader disagreed, so the author conceded" is precisely the circularity the engine refuses to certify — and this repo's own rule is that *an instrument with no disagreement left in it measures nothing* (`g_054` has been deliberately contested since v0.7 for the same reason). The author's readings are defensible; so are the grader's; the disagreement is now data rather than noise, and it is disclosed rather than laundered. What *was* corrected — `g_076`, once — was corrected because all three runs agreed **and** the author's reading failed on its own terms (crediting an intermediate of the wrong computation). One concession on evidence is adjudication; three is capitulation. --- ## What the badge says now, and why it changed **Before:** `0 of 198 · it has never once inflated a grade` — 66 items × 3 runs. **That claim did not survive the extended set.** Across three audits (774 judgments) this grader inflated **3 times**, and every one was traced to an ambiguity in the *spec*, not to the model being lenient by nature: 1. `cap at partial` read as an instruction to **award** partial rather than to limit it. 2. A right-answer-wrong-method tiebreak added *by this release* generalized into a universal criteria-counting rule that overrode `lapsed = core absent or wrong`. Both are now closed, and the shipping spec measures **0 inflations in 258 judgments**. The badge therefore reads **0 / 258**, and the README states plainly that the number was re-earned after three inflations were found and fixed — because a safety badge that has never been stress-tested is worth less than one that was, failed, and was repaired.