Freeze point: protocol as of commit b12c6fc (2026-08-03). Every deviation from protocol-specified method is logged here with a reason.
| # | Date | Deviation | Reason | Effect on findings |
|---|---|---|---|---|
| 1 | 2026-08-03 | H2.1 adjudicated from published NSECE snapshots instead of SIPP/NSECE microdata | Microdata acquisition/processing exceeds session scope; published ACF snapshots are the same instrument's official tabulations | Precision reduced (±pp unknown); direction unaffected; microdata pass remains queued |
| 2 | 2026-08-03 | Scorecard double-scoring done as same-analyst shuffled re-score in the same session, not two analysts or a time gap | Single-agent execution | Weaker independence than §14.2 intends; discharged 2026-08-07 by structurally blinded full-matrix re-score (deviation #10) |
| 3 | 2026-08-03 | §16 sequencing drafted (v0) before the formal scorecard ran | Pass-1 synthesis value | Sequencing to be re-checked against formal scorecard output this pass; discrepancies noted in ws16 |
| 4 | 2026-08-03 | FTE model uses representative licensing ratios and coverage scenarios rather than state-by-state ratio tables | Session scope; parameters are explicit and adjustable in the committed script | National totals robust to ±1 ratio point (sensitivity shown); state detail queued |
| 5 | 2026-08-03 | Scorecard consolidated to 14 dimensions from the protocol's 23 | Omitted dimensions fold into retained ones or lack pass-level evidence; scored narratively in rationale where evidence exists | Full-dimension pass queued with second-reviewer re-score |
| 6 | 2026-08-03 | Re-score performed on a 3-architecture sample, not the full matrix | Session scope | Two ±1 adjustments applied; top-4 unchanged; discharged 2026-08-07 by full-matrix re-score (deviation #10) |
| 7 | 2026-08-03 | Anchor rows 6, 7, 10, 13–18, 20 remain unverified after pass 2 | Search returns did not cover them; two-source rule enforced (no fill from memory) | Rows blank, not guessed; queued for pass 3 |
| 8 | 2026-08-03 | Pass 3 closed anchors 6, 10, 14, 15; rows 7 (initiative states), 16 (background-check backlogs), 20 (Head Start/Lanham dates) remain blank | Not returned by searches; two-source rule enforced | Blank, queued; none is load-bearing for current conclusions |
| 9 | 2026-08-03 | ARPA price-effect replication attempted and found absent — OpenAlex full-corpus check returned 4 items, none evaluative | The literature does not yet exist, not a search failure | ws09 conclusion holds at single-source-family confidence; register entry added to §17 |
| 10 | 2026-08-07 | Independent structurally-blinded re-score of full 12×14 matrix (batch API; evidence inlined; scorecard unreachable) | M6/M10 owed after #2/#6 | 47 cells corrected, 25 judgment splits logged; stable set a3/a4/a8; a9 exits; a5+a10 last under every weighting; see scorecard/rescore-log.md |
| 11 | 2026-08-08 | Verification Protocol Phases 1–2 artifacts landed from PR #38 salvage (ledgers, verification-log, steelman suite, method/sources atlas); site status/hero edits from #38 discarded as stale post-#66 | #38 was CONFLICTING with main after independent re-score; triage kept research record, dropped regressive publication chrome | Evidence-tag claim on scales.md corrected; full claim ledger + steelmen now in record; public-page fact corrections from #38 still owed as a follow-up pass against current site text |
| 12 | 2026-08-10 | Verification Protocol Phase 1 CORRECTED/OVERSTATED items landed against current (post-#66) public pages: Part 4 Byrd-rule rewrite (A1–A7, A9, A10); DOD/Australia/Quebec/Georgia-Florida precedent-table corrections (G1, G3, G5, G8); workforce baseline, turnover, credential-pipeline, ARPA access-rate, and wage-parity-denominator figures (B5, B9, F1, F2, F4, F8, F9, C1, C2, G7); Bright Horizons center count and Tri-Share tenure (B15b, H4); rural-desert and CDFI framing (H2, H3); CCDF final-rule rescission (I1); architecture count corrected to 11 + cash comparator (D6); scorecard-scales evidence-tag note corrected on its live mirror page (D1) | Deviation #11 flagged these as still owed against current site text | site/childcare/index.html, site/childcare/sources/index.html, and site/childcare/sources/scorecard-scales/index.html now match verification-log.md. Left unedited (STALE/UNVERIFIABLE/REFER-UP, out of this pass's scope): B1, B4a, B15a population/wage/occupancy staleness; A12 CBO 57630; H1 220,000-provider citation; F5 DC PEF FY-year label (ws04-workforce.md, mirror page only); F7 Treasury attribution; G2 Germany courts (explicitly REFER-UP in the log, not rewritten) |
| 13 | 2026-08-10 | Verification Protocol Phase 2 (steelman, steelman-log.md, built 2026-08-04) ran in the same session as Phase 1's fact-check, deviating from S1's requirement that Phase 2 run in a session independent of the checker's conclusions |
Run at the user's instruction; mitigation used instead — each of the three steelman targets (A: cash/employer instruments, B: regulatory constraint, C: is universal 0–2 care desirable) was built in a fresh agent context that received the corrected filing but not the checker's reasoning about where it was weak, per S2 | Logged here per §19, not only in steelman-log.md's own preamble, so a reader who considers session-independence necessary can discount accordingly. The structural findings that survived (scorecard dimensions anchored on instrument type, the child-outcome term missing from child_dev_first, the unscored regulatory alternative) are independently checkable against the committed record regardless of this deviation |
| 14 | 2026-08-10 | Verification Protocol Phase 2 (steelman) findings landed on the public whitepaper and sources digest: Part 2's "Binding constraint" stamp qualified (ratio relaxation to Florida's own statutory levels matches halving turnover's hiring reduction, at no wage cost); Part 6's "three designs survive every ranking" qualified (three of fourteen scorecard dimensions anchor their scale on instrument type rather than outcome — dropping them changes the top four under all four weightings); honesty box extended with the Head Start Impact Study omission, child_dev_first's missing child-outcome term, and Quebec's outcome literature being cited only for its waitlist |
Corrections are visible, never silent (protocol §4); Phase 2's S4/S5 verdicts are consequences owed to the record, not footnotes | site/childcare/index.html and site/childcare/sources/index.html now disclose where the filing's own instrument and scorecard structure foreclose questions it presents as tested. Corrected 2026-08-11 (deviation #15): this entry originally recorded that a3 (Head Start scaled to universal) "leads under every weighting and every sensitivity tested." It does not, and did not on the day this was written — a8 (public option) leads all four weightings on the committed board |
| 15 | 2026-08-11 | Steelman A's rank-sensitivity table (steelman-log.md A-1) and every claim derived from it were computed against the pre-2026-08-07 board and never recomputed after deviation #10's blind re-score moved 47 cells. The stale table was committed 2026-08-08 and published to the whitepaper, sources digest and corrections log on 2026-08-10. Its central derived claim — "a3 leads under every weighting and every sensitivity tested" — is false on the committed rankings, where a8 leads all four |
Caught on review against scorecard/rankings.txt, which has been unchanged since the re-score. Nothing new was discovered: the claim was true of the superseded board and was carried forward without re-running it |
Table recomputed and corrected in place with the supersession noted; a3-leads claims corrected in steelman-log.md, verification-log.md, the whitepaper Part 6 and honesty box, the sources digest, and the gbmt-1-scorecard-instrument corrections entry. scorecard/rank-sensitivity.py committed — it reproduces the committed rankings and aborts before computing any variant if they do not match, so this specific failure cannot recur silently. The structural finding survives and is sharpened: dropping the three instrument-labelled dimensions changes the top four under all four weightings, and puts the cash comparator into the equity-first top four |