Protocol: method/verification-protocol.md, Phase 1 (fact-check). Filing under check: Whitepaper No. 1, filed 2026-08-03 — site/childcare/index.html, site/childcare/sources/index.html, and the research record in childcare/docs/ and childcare/research/. Check date: 2026-08-04. Independence (V1): this session did not author GBMT-1 and gave the filing's own red team no evidentiary weight. Source atlas: primary texts fetched for this pass are committed under method/sources/ with provenance headers.
Coverage (V3)
| Ledger | Rows |
|---|---|
Public pages (site/childcare/index.html, site/childcare/sources/index.html) |
326 (P001–P326) |
Research record (childcare/docs/, childcare/research/, childcare/baseline/) |
522 (R001–R522) |
| Inferential repair pass (see "Extraction guardrail") | 181 (I001–I181) |
| Total extracted | 1,029 |
Coverage proof: 1,029 extracted · 1,029 verdicts · 0 unaddressed. The clusters below account for every row. Two honest caveats on that arithmetic, stated here rather than at the foot where they would be easier to miss: a substantial minority of verdicts are CONFIRMED-INTERNAL (reproducible from the repo, no external referent) or confirmed at secondary tier because 403s blocked the primary source, and ~68 rows are UNVERIFIABLE. An unverifiable verdict is a verdict, not a verification. Every claim was reached and given a disposition; not every claim was reached at the tier V2 asks for.
Extraction guardrail (V3) — the delegated ledger failed its own test
The protocol requires the adjudicating model to independently re-extract one section and diff it against the delegated ledger before trusting it. That guardrail fired.
Sample section: Part 4, "Senate procedure: finance survives, mandates die" (site/childcare/index.html:192-198).
- Independent re-extraction by the adjudicating model, recorded before reading the delegated ledger: 12 substantive assertions (plus 3 source attributions).
- Delegated ledger for the same lines: 4 rows (P025–P028).
- Miss rate: 8 of 12 (67%).
The misses were not random. Every dropped row was a causal claim, statutory characterization, or policy inference:
| Dropped assertion | Type |
|---|---|
| "the controlling precedent" (the precedential status itself) | CHARACTERIZATION |
| "struck ... as a private-sector mandate" | STATUTORY |
| "A childcare wage mandate dies identically" | INFERENCE |
| "wage funding is money and survives" | STATUTORY |
| "Quality standards are safest housed inside federal spending programs" | INFERENCE |
| "the holdout problem is the official base case" | CHARACTERIZATION |
| "any national design requires a federal backstop" | INFERENCE |
The delegated pass captured every number and date in the section and dropped the reasoning. That is the exact material V4 calls the INTERPRETATION failure mode — the mode the protocol was written to catch. An extraction layer with this bias does not merely under-count; it removes the protocol's subject matter from the ledger before adjudication starts.
Action taken: the extraction layer was repaired by re-running it at a higher tier against an explicit inferential-claim taxonomy, and re-tested against the same Part 4 guardrail. Consequence for the orchestration design: logged as a correction to method/verification-protocol.md — see the deviations entry.
Cluster A — Reconciliation and the Byrd rule (the protocol's named highest-risk claim)
The protocol singled this out: "its load-bearing procedural analysis (the Byrd rule, reconciliation) is the kind of technical legal claim that fails quietly." It did fail — in a more specific way than expected, and in one respect the filing was closer to right than the secondary record suggests.
Filing's sourcing tier. childcare/research/ws11-legislative.md:7 rests this analysis on a Manhattan Institute article and an EPIC for America explainer — both advocacy-affiliated secondary analyses. The same file records "CRS deep-dive pending" and lists "CRS Byrd compilations" under Pending. The claim was published on the public whitepaper as "the controlling precedent" while the record itself marked its verification as outstanding. That gap is the finding behind all of the verdicts below.
Primary sources reached this pass: 2 U.S.C. § 644 (statute); CRS RL30862, The Budget Reconciliation Process: The Senate's "Byrd Rule" (Sept. 28, 2022), including its table of Byrd rule points of order; CRS R48640, The Senate's Byrd Rule: FAQ (Aug. 21, 2025). All committed to method/sources/.
What the primary record actually shows
CRS RL30862's table of Byrd rule points of order, entry 22 — American Rescue Plan Act of 2021 (P.L. 117-2), records:
- "To Strike Provision(s) from Bill or Conference Report: [none]"
- "To Bar Consideration of Amendment(s): Sanders Amendment No. 972 — basis: "Budgetary changes merely incidental to non-budgetary components" — subject: "To provide for increases in the Federal minimum wage" — waiver motion Rejected, 42-58 (March 5, 2021) — disposition: "Sustained; amendment fell."
Two distinct events, which the filing merges into one:
- Feb. 25, 2021 — the Parliamentarian advised that the House-passed minimum-wage provision did not comply; Senators removed it from the substitute before floor consideration. Per CRS R48640: "When Senators respond to guidance from the Senate Parliamentarian by choosing to remove or redraft questionable matter in a reconciliation measure prior to its consideration on the Senate floor, then no formal precedent is created, and information that may guide future interpretations of the rule is not published."
- March 5, 2021 — Sanders Amendment No. 972 sought to add it back; a Byrd point of order was raised on merely-incidental grounds, the waiver failed 42-58 (60 required), and the point of order was sustained.
So nothing was struck from the Act; an amendment to add it was barred. But a Byrd determination on merely-incidental grounds was sustained by the Chair and is recorded in CRS's compilation — which is a stronger footing than the pre-floor-removal story the secondary sources tell.
Verdicts
| # | Claim | Location | Verdict | Basis |
|---|---|---|---|---|
| A1 | "Reconciliation's Byrd rule strips provisions whose budgetary effect is 'merely incidental'" | index.html:195 | OVERSTATED | Accurate for § 644(b)(1)(D) alone. The Byrd rule defines six categories of extraneous matter, (b)(1)(A)–(F). Presenting one test as "the Byrd rule" is what makes A5 possible. |
| A2 | "the controlling precedent being the $15 minimum wage" | index.html:195 | OVERSTATED | A sustained point of order exists, so this is not baseless. But CRS: the merely-incidental test "requires an exercise in judgment" and "every determination is necessarily case-specific, complicating efforts to inform a bright line." "Controlling" asserts a bright line CRS says does not exist. |
| A3 | "struck from the 2021 relief act" | index.html:195 | CORRECTED | CRS RL30862 records "[none]" struck from P.L. 117-2. The provision was dropped pre-floor after Parliamentarian advice; separately, an amendment to add it was barred by a sustained point of order. Nothing was struck from the Act. |
| A4 | "as a private-sector mandate" | index.html:195 | CORRECTED | The recorded basis is "budgetary changes merely incidental to non-budgetary components." "Private-sector mandate" is not a category in § 644 and is not the recorded ground. The sources page states this correctly ("on grounds its budgetary effect was 'merely incidental'") — the error was introduced when the record was translated to the public whitepaper. |
| A5 | "anything that moves money survives; anything that mandates behavior dies" | sources/index.html:224; ws11-legislative.md:10 | CORRECTED | Directly contradicted by CRS R48640: "a provision with significant fiscal effects but even larger policy effects could be deemed extraneous. Conversely, a provision with small budgetary effects and negligible policy effects could be permissible." Independently refuted by § 644(b)(1)(E), under which a plainly money-moving provision is extraneous if it increases outyear deficits — the very rule the filing relies on elsewhere when it says BBB "was written time-limited to fit the budget window" (ws11:13) and that reconciliation sunsets are "a design defect" (index.html:230). The filing's generalization contradicts its own adjacent finding. |
| A6 | "A childcare wage mandate dies identically" | index.html:195 | OVERSTATED | Direction is supported; "identically" is not, given case-specificity. Replace with "faces the same objection." |
| A7 | "wage funding is money and survives" | index.html:195 | OVERSTATED | Survives (b)(1)(D), which is the filing's point and is sound. But "survives" unqualified ignores (b)(1)(C) jurisdiction and (b)(1)(E) outyear-deficit exposure — the latter being why the 2021 title was written time-limited. |
| A8 | "Quality standards are safest housed inside federal spending programs rather than imposed on the market" | index.html:195 | CONFIRMED (as comparative) | Consistent with the (b)(1)(D) comparison: conditions on new spending have a larger budgetary component relative to non-budgetary than market-wide mandates. Stated comparatively ("safest"), not absolutely — correctly hedged. |
| A9 | "Quality provisions are the most Byrd-vulnerable part of any bill" | sources/index.html:205 | OVERSTATED | "Most ... of any bill" is a superlative the record does not establish; no ranking against other provision types was performed. |
| A10 | Filing's characterization of the Parliamentarian as having "removed" the provision | sources/index.html:224 | OVERSTATED | The Parliamentarian advises; the presiding officer rules; Senators removed it. § 644(a) requires a point of order "sustained by the Chair" to strike. Sanders' own Senate statement (Feb. 25, 2021) is titled "Statement on Parliamentarian's Advice." |
| A12 | "CBO's scoring of the House-passed BBB title ($381.5B, 2022–31) embedded substantial state non-participation" | index.html:195; sources anchor 1 | UNVERIFIABLE | CBO publication #57630 could not be obtained in primary form. cbo.gov serves a DataDome bot challenge to every automated route (browser UA, multiple proxies, direct PDF); the Wayback Machine's availability API returns {"archived_snapshots": {}} for both the publication page and the PDF — no snapshot exists to recover. Every located third-party reference resolves back to cbo.gov. Full attempt table in method/sources/UNOBTAINED.md. Per V5 this is UNVERIFIABLE, not CONFIRMED: the $381.5B figure and the state-non-participation assumption are secondary-reported only. This is a load-bearing claim — it carries anchor 1 and the §6 federal-backstop conclusion — so the filing should carry an explicit "secondary-sourced" flag on it until the document is reached. |
| A13 | The Parliamentarian's specific reasoning on Feb. 25, 2021 | ws11; sources/index.html:224 | UNVERIFIABLE | No written ruling exists publicly, and this is structural rather than a retrieval gap: Byrd-rule point-of-order advice is delivered verbally to the presiding officer and not published. That a determination occurred is verifiable at agency-primary tier (Senate Roll Call Vote 74, 117th Cong. 1st Sess., March 5 2021, waiver failed 42-58; Sanders' official statement of Feb. 25 using "advised"). Which subsection the Parliamentarian invoked on the earlier date rests on unnamed-source press accounts only. This does not disturb A1–A4: the "merely incidental" ground for the sustained point of order comes from CRS RL30862's own table, not from press. |
| A11 | Implied negative: no minimum-wage provision reached the enacted law | — | CONFIRMED (V5, positively verified) | Checked against the enacted text itself, not a summary: P.L. 117-2 fetched from govinfo.gov (243-page official PDF) and searched in full. Zero occurrences of "minimum wage" across 748,902 characters; zero references to 29 U.S.C. 206 (the FLSA minimum-wage rate); the only two Fair Labor Standards Act references are definitional, in the Title IV aviation-manufacturing provisions. Text committed to method/sources/. |
Replacement wording (public whitepaper, Part 4)
Reconciliation's Byrd rule declares six categories of provision extraneous; the one that governs here strips provisions whose budgetary effects are "merely incidental" to their non-budgetary components. The closest on-point determination is the $15 minimum wage: the Parliamentarian advised in February 2021 that it did not qualify, it was dropped from the American Rescue Plan before floor consideration, and when it was offered again as an amendment that March a point of order was sustained against it on merely-incidental grounds, the waiver failing 42-58. A childcare wage mandate faces the same objection; wage funding is budgetary and survives that test — though CRS cautions every such determination is case-specific, and money-moving provisions can still fall to the rule's outyear-deficit test, which is why reconciliation-built programs get written with sunsets.
What this does and does not change
It does not overturn the filing's operative conclusion. "Fund compensation rather than mandate it, and house standards inside federal spending programs" survives — the primary record supports the direction. What falls is the confidence and the mechanism: a case-specific judgment was published as a bright-line rule ("dies identically," "controlling precedent"), and a compressed generalization ("anything that moves money survives") contradicts both CRS and the filing's own sunset finding. Per the protocol, this is applied to the record, not footnoted.
Cluster B — bounded numeric and citation claims
Delegated to the low-risk adjudication layer per the efficiency architecture, with primary-source verification required and search-summary figures barred. 22 claims checked: 12 CONFIRMED · 3 CONFIRMED-with-caveat · 4 CORRECTED · 2 STALE · 1 negative positively verified.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| B9 | "a jump in parents unable to find care, 17.7% to 22.2%" | CORRECTED | The figures appear nowhere in the cited CEA source. Re-verified independently by the adjudicating model: fetching the CEA brief and stripping SVG markup leaves zero occurrences of "17.7" or "22.2" in the prose. The source says, verbatim: "Figure 6 shows that across the country, the share of households who could not access care increased from 24% to 31% after the expiration of ARP funds (between Q3:2023 and Q1:2024)." The digits do occur in the raw HTML — inside unrelated SVG icon path data. This is the fabricated-figure failure mode the protocol's traps section names, and it reached a live public page. Note the direction: the real jump is 7pp, not 4.5pp — the correction makes the filing's finding stronger, not weaker. |
| B5 | "against ~1.05 million today" | CORRECTED | fte_model.py:41 labels the constant in its own comment as "industry jobs, pre-pandemic baseline" — a Jan 2020 figure (BLS CES6562440001 = 1,045,800). Current employment is ~1.10M. The word "today" is wrong, and the error inflates the stated shortfall. |
| B15b | "Bright Horizons closed 90 centers in 2025" | CORRECTED | The 10-Q states a net reduction of 90 centers cumulatively between Dec 2022 and March 2026 (~3.25 years). 2025-specific guidance was a net 25–30. The record (ws12-private-capital.md:8) had both numbers side by side; the public page collapsed them into the wrong one — the translation-drift pattern again. |
| B15a | "chain occupancy ~71%" | STALE at the filing date | True for Q2 2025 only. Q3 2025 (67%) and Q4 2025 (64.5%) were both public before the 2026-08-03 filing. Not a supersession-since-filing — it was already out of date when published, so this is a sourcing miss rather than a V6 STALE. |
| B1 | Population bands (51.2M / 11.1M / 7.5M / 32.6M) | CONFIRMED, now STALE | Reproduced exactly from the raw Census Vintage 2024 file. But Vintage 2025 (released 2026-06-25, five weeks before filing) gives 50.6M / 11.0M / 7.5M / 32.1M. |
| B4a | "$15.41/hour" | CONFIRMED, now STALE | Verbatim on the BLS OOH page for May 2024. The May 2025 OEWS release (2026-05-15, before filing) supersedes it; BLS's own summary page had not yet been refreshed. |
| B11 | Negative: "no peer-reviewed evaluation of the $24B experiment exists" | CONFIRMED (V5) | Independently searched Crossref, NBER, arXiv and general web; no evaluation found. The filing's specific "four papers" OpenAlex count could not be reproduced (API rate-limited) — the claim holds, the count is unverified. |
| B20 | 1971 CCDA vetoed "December 10, 1971" | CONFIRMED-with-caveat | The veto message is dated Dec 9, 1971 and was transmitted to the Senate Dec 10. Both dates are defensible; the filing should say which it means. |
| B10, B13, B17, B18, B19, B21, B22 | Australia +4.4%; Canada 194K/284K; 17 supermajority states; FSA cap; GA/OK/FL dates; Netherlands scandal; Tennessee/Boston studies | CONFIRMED at secondary tier | Each confirmed, but 403s blocked the primary agency pages, so these sit one tier below V2's standard. Listed so the gap is visible rather than implied closed. |
Cluster C — the committed FTE model (adjudicated directly)
The whitepaper's headline finding — the workforce is the binding constraint — rests on a model in this repo rather than on an external source, so it was re-run rather than looked up.
childcare/baseline/scripts/fte_model.py reproduces every published figure exactly:
| Published claim | Model output | Verdict |
|---|---|---|
| "~2.8 million educator FTEs" | central/central: 2,807k | CONFIRMED |
| "range 1.9–3.9M" | min 1,915k (low/loose), max 3,941k (high/tight) | CONFIRMED |
| "~800,000 hires per year" | 798k | CONFIRMED |
| "579,000 merely replace departures" | 579k | CONFIRMED |
| "cutting turnover in half saves more hiring than the entire expansion requires" | halving 0.30 turnover saves ~289k/yr vs. 220k/yr net growth | CONFIRMED |
| "$15.41/hour" median | WAGE_NOW = 32,050 ÷ 2,080 h = $15.41 |
CONFIRMED (internally consistent; the BLS figure itself sits with Cluster B) |
This is the committed-pipeline discipline working as advertised: the arithmetic is reproducible from the repo by a checker who did not write it. Two qualifications:
| # | Claim | Verdict | Basis |
|---|---|---|---|
| C1 | "wage parity is the largest single cost term (~40%)" (sources/index.html:190) | OVERSTATED | The model prints two different ratios for this quantity — 22% of gross parity cost, 65% of the parity-vs-current-wage increment — and the ~40% comes from a third denominator introduced in ws04-workforce.md:27 (a ~$42B private-spend baseline). All three are defensible; the public page cites one without saying which. This series has already made denominator-dependence a headline finding once, in elder care's caregiving-valuation row. Same defect, this filing. Replacement: "~40% of incremental cost measured against current private spending — the figure moves with the denominator." |
| C2 | H4.1 presented as a finding on the public page | OVERSTATED | ws04-workforce.md:27 adjudicates H4.1 INDETERMINATE ("below the 50% support threshold, above the 35% refutation floor"). The sources page states the ~40% conclusion without the indeterminate label. The record is more careful than the page again — the same direction of drift as Cluster A. |
Model caveats already disclosed by the filing and confirmed as disclosed: representative rather than state-by-state licensing ratios (deviation #4), and a CURRENT_WORKFORCE constant of 1.05M carried as a parameter rather than a live pull.
Cluster D — the scorecard (the pass's most serious finding)
The scorecard produces the filing's central architecture recommendation, and the series method (M6) requires "anchored ordinal scales authored first; every cell cites a workstream finding; rankings under multiple objective weightings with rank-stability reported; independent re-score with a reconciliation log." Audited cell by cell, then re-run.
What reproduces. rank.py regenerates rankings.txt exactly (diff empty). The matrix really is 14 dimensions × 4 weightings. The reconciliation log's shuffled sample re-score is real and its two applied adjustments are recorded. That part of the method was delivered.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| D1 | scales.md: "Every cell in the CSV carries an evidence tag resolvable in the rationale file" — published live at site/childcare/sources/scorecard-scales/ |
CORRECTED | scores.csv holds 168 cells (12 rows × 14 dimensions). The rationale's "Key cell citations" section tags 9. The claim is false by two orders of magnitude, and it is the filing's headline methodological assertion about its own central instrument. Verified by direct count of both files. |
| D2 | Method M6: "every cell cites a workstream finding" | NOT MET for this filing | Same evidence. Recorded here rather than silently, because M6 is the series standard and every other filing's scorecard was built the same way. This should be checked in the other five filings before their scorecards are cited again. |
| D3 | rationale.md: a5 integrity_risk=1 cited to ws15 for the Netherlands clawback |
CORRECTED (citation) | ws15-precedents.md contains zero occurrences of "Netherlands" (verified by search). The material lives in the anchor table (row 14) and protocol-review.md. The filing's own red team, Attack 1, asked for it to be added to §15 — that never happened, so the tag points at a document that does not contain the fact. The underlying fact is sound; this is a citation repair, not a withdrawal. |
| D4 | rationale.md: a7 durability=5 cited to ws04's DC-PEF passage |
OVERSTATED (citation) | The cited passage documents DC-PEF lacking dedicated revenue and fighting annually to survive — evidence for scoring a different architecture down, not for this scale's top anchor ("broad constituency AND dedicated revenue AND institutional entanglement"). The score may be right; the citation does not carry it. |
| D5 | rationale.md + sources page §14: "Top 4 under all four objective weightings, no flips" |
CORRECTED | rankings.txt contains two flips inside the top four: a3/a8 swap under labor_supply_first (exact tie, 3.85 each) and a6/a9 swap under equity_first (exact tie, 3.65 each). The set is stable; the ordering is not, and both flips are ties where the printed sequence is arbitrary. |
| D6 | Whitepaper: "Twelve candidate architectures were scored"; sources page: "12 candidate architectures plus a cash comparator" | CORRECTED | Neither is right, and they contradict each other. scores.csv has 12 data rows: a1–a11 (eleven architectures) plus cash_comparator. The whitepaper counts the comparator as an architecture; the sources page counts twelve and then adds the comparator again. |
| D7 | "Anchored scales authored before scoring" (scales.md title; method M6) |
UNVERIFIABLE | scales.md, scores.csv, rationale.md, rank.py and rankings.txt were all added in a single commit (efdce2ae, 2026-08-03). Git preserves no intra-commit ordering, so the record cannot demonstrate the sequence. The claim rests on stated intent. Not an accusation — a note that the filing's own committed record does not evidence a process claim it makes. |
S5 — rank sensitivity under Phase 1's corrections
The corrected ws11 finding leaves one comparative standing (A8, CONFIRMED): standards housed inside federal spending programs are the most Byrd-defensible form. That bears on procedural for the two architectures that do exactly that — a3 (scored 3) and a8 (scored 2). Re-scoring both to 4:
The top-4 set does not move under any weighting — the filing's recommendation survives. What moves is a3's primacy: a8 ties a3 under child_dev_first (3.95 each) and leads under labor_supply_first (3.95 vs 3.90). scores.csv is left at committed values — the re-score is a reported sensitivity, not a silent restatement.
Superseded 2026-08-11. This sensitivity was computed against the board as it stood in Phase 1, before deviation #10's structurally-blinded re-score of 2026-08-07 moved 47 cells. It is left here as the dated Phase 1 record, but it no longer describes the committed board: a8 now leads all four weightings outright, so there is no a3 primacy for the procedural re-score to erode. The full table this paragraph cited to
scorecard/rationale.mdwas replaced when that file was rewritten for the re-score; the current sensitivity of record isscorecard/rank-sensitivity.py. See deviation #15.
Why D1 is the pass's most serious finding. The Byrd corrections changed how confidently one sentence could be stated. D1 changes what the scorecard is: not an evidence-linked instrument but a single analyst's judgement applied against scales they wrote, with ~5% of cells tagged. That is a legitimate method — but it is the method the filing's own red team, Attack 5, warned about in its own words ("if the scores encode the priors, stability across weightings measures internal consistency, not truth"), and the record claimed a stronger one. The honesty box already says "one analyst, one desk" and flags the missing second scorer; it does not say the cells are largely uncited, and scales.md says the opposite.
Cluster E — the meta-record (274 rows, all adjudicated)
protocol-review.md (97) · research-inquiry.md (21) · adjudication-criteria.md (22) · deviations-log.md (27) · red-team.md (43) · steelman-feasibility.md (33) · incompatibility-log.md (32). 274 of 274 adjudicated, zero unaddressed in this batch.
Tally: ~178 CONFIRMED-INTERNAL · ~23 CONFIRMED external · ~68 UNVERIFIABLE · 2 OVERSTATED/CORRECTED · 3 transcription slips.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| E1 | "Pre-registered adjudication criteria written before evidence" (sources page receipts; method M3) | UNVERIFIABLE — and the public wording overstates what the record can show | adjudication-criteria.md was added in commit efdce2a, whose own title is "Pass 2: adjudication criteria, FTE model, formal scorecard, red team" — the criteria, the model they adjudicate, the scorecard and the red team all landed together. Git orders only at commit granularity, so criteria-before-evidence is asserted, not shown. Same shape as D7. Not an accusation of bad faith: pass-1 findings did land earlier. But "written before evidence" is a verifiable-sounding claim the committed record cannot verify, and it is printed on the public page. |
| E2 | steelman-feasibility.md: "Two of the three are Medicaid holdout states" |
OVERSTATED at the filing date (fact since confirmed) | ws06-delivery.md:11 logs the holdout count as "pending, flagged for verification (two-source rule not yet met)." The steelman states it as settled. The caveat did not travel with the claim. Independently checked this pass: the number does hold — GA and FL are non-expansion, OK expanded in 2021. So the filing was right, but it did not know it was right when it published. This is the translation-drift pattern inside the record itself, not just from record to public page. |
| E3 | incompatibility-log.md #6: the "state-funded only" caveat travels with every NIEER figure |
CORRECTED | Neither NIEER-sourced figure in baseline-v0.md carries the caveat the log says accompanies them. The log describes a practice that was not implemented. |
| E4 | baseline-v0.md: PR/territory fix queued in "Next Increments" |
CORRECTED | Not actually present in that section. |
| E5 | R063, R066, R084 | Ledger artefact, not a filing error | The source documents mark these "(verify)"; the extraction pass lifted them as settled claims. The filing hedged correctly; the ledger dropped the hedge. Worth noting because it is the extractor introducing false confidence — a third variant of the same failure. |
| E6 | The red team's five attacks | CONFIRMED | All five stated dispositions match the evidence they cite. Attack 4's "amendment applied" was checked directly against ws16-sequencing-v0.md, which carries the exact restated language. The red team did what it said it did — worth stating plainly given how much else in this cluster is qualified. |
| E7 | ~68 external claims embedded in protocol-review.md and steelman-feasibility.md prose |
UNVERIFIABLE | UK/Netherlands/NYC precedent details, academic attributions, GAO/OIG references appearing inside self-admittedly recollection-based review prose. No primary source reached this pass. These are not load-bearing for any published conclusion, but they are unverified and now marked as such rather than carried as background fact. |
Coverage gap noted against this pass itself: deviations-log entries #10–#16 — this verification pass's own self-documentation — post-date the extraction ledgers and appear in none of them. Recorded here so the coverage arithmetic below is not quietly wrong.
Cluster F — workforce, cost and baseline (171 rows)
~110 CONFIRMED (including CONFIRMED-INTERNAL, reproduced from committed data and scripts) · 5 CORRECTED findings across ~25 rows · 3 OVERSTATED across ~5 · 2 UNVERIFIABLE across ~7.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| F1 | "10th-lowest-paid of 825 occupations" (whitepaper abstract + Part 2 + sources §4) | UNVERIFIABLE, with contrary indication | BLS blocks automated access from this environment, so the rank could not be verified at source. CSCCE's own index implies a materially different position (~top-3% lowest, nearer 20th–25th). A precise-sounding rank that cannot be reached, and whose nearest available check disagrees, should not be stated as fact. Replaced on both public pages with "among the lowest-paid occupations BLS tracks" — which is certainly true and is all the argument needs. |
| F2 | "The credential pipeline issues ~40,000 per year, renewals included" / "under a fifth of growth hiring" | CORRECTED | The cited CDA Council URL contains no figure at all. The Council's own 2024 annual report — published before this filing — reports a record >50,000/yr, making the pipeline ~23% of net-new hiring, not "under a fifth." The conclusion survives: >50k still cannot carry a ~798k/yr ramp. Corrected on the public page. |
| F3 | Warren-style plans "$600B+" (10-yr) | CORRECTED | The cited Cornell source states **70B/year * *andgivesnoten − yeartotal.Naive × 10is 700B. The filing's own figure is not in its own source. |
| F4 | Cleveland Fed "turnover ~65% above a typical job, 2010–2022" | OVERSTATED | The source ties 65% to 2022 specifically; the period average is ~64%. Direction holds, the framing as a period average does not. |
| F5 | DC Pay Equity Fund "$70M FY26 vs $92.4M FY27" | CORRECTED | $70M was FY2025; the elimination fight and $92.4M ask were FY2027. DC Council has since passed $73.5M for FY2027 (pending Congress) — unreported in the record. |
| F6 | FCC-home decline / center growth figures (R017–R018) | CONFIRMED, wrong citation | Figures are right; the linked Urban Institute PDF does not contain them. The source is CCAoA's Catalyzing Growth 2022. Citation repair, not withdrawal. |
| F7 | "$42B private spend," attributed to Treasury | UNVERIFIABLE | Corroborated at secondary tier (Fed Communities); the Treasury attribution specifically could not be reached. This figure is the denominator behind C1's "~40%" wage-parity share. |
| F8 | "Every published cost estimate excludes school-age" | OVERSTATED | The record's basis is four collected estimates, not an exhaustive survey. "Every estimate we collected" is supportable; "every published estimate" is not. |
| F9 | "Labor share swamps everything else" | OVERSTATED | The 50–80% labor share is confirmed; no facility/admin breakdown exists in the record to support "swamps." |
Cluster G — precedents, quality and politics (74 rows)
51 CONFIRMED · 9 CORRECTED · 9 OVERSTATED · 2 UNVERIFIABLE · 1 REFER-UP.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| G1 | "US DOD — pay parity + standards — works; has for decades" (Part 5 table, stamped Verified) | CORRECTED — the pass's second-most serious finding | No citation for this existed anywhere in the record, and the filing's own to-do list still carried "DoD deep-dive" as pending. The statute, 10 U.S.C. § 1792(c), sets a "competitive local-market rate" — not parity with any named benchmark. GAO-24-106524 (May 2024) found 34–50% turnover in FY2022 and staffing-driven waitlists across all four services. Read the consequence carefully: DoD is a fifth instance of funding outrunning staffing, which strengthens the filing's central workforce finding — while destroying the one example it offered that competitive pay had ever been tried at scale. The honesty box's "Nobody has tried fair wages at scale" becomes more true, not less. |
| G2 | Germany: "courts award money, not slots" | CORRECTED / REFER-UP | Two tracks exist: BGH (2016) awards lost-wage damages, fault-dependent; VGH Baden-Württemberg (2022) orders actual placement by injunction and rejects staff shortage as a defence. Courts compel both money and (nominal) slots. Flagged for legal-specialist review rather than rewritten here — foreign-law interpretation is above this pass's tier. |
| G3 | Australia: "strings later attached" | CORRECTED | The ACCC's own December 2023 report found the hourly rate cap had "only limited effectiveness… does not act as an effective signal of high prices." The record framed this as a fix that worked; the primary regulator says it did not. Corrected on the public pages. |
| G4 | Australia "+4.4% in a representative year" | OVERSTATED | The figure is real (one quarter's YoY, via a Guardian report of a DoE release). "Representative year" is the filing's own unsourced gloss. |
| G5 | Quebec: "staffing named as cause" | OVERSTATED | The cited CBC article does quote a CPE director naming staffing — and the same article quotes unions and opposition naming inadequate funding. The two causes are linked, not competing; the filing presented only the half that fit. |
| G6 | Quebec ">60,000 shortage" | CORRECTED (citation) | Figure is real but absent from the cited URL; it comes from a different December 2024 article. True but unlanded. |
| G7 | DC PEF "turnover measurably fell" | OVERSTATED | The primary study (Urban Institute) shows a cross-sectional 37% vs 51% gap among subsidy-accepting centers, marginally significant at p=.07 — not a measured before/after decline. This is load-bearing: it is the whitepaper's evidence that paying more fixes retention. Corrected in the Part 2 plain-talk box. |
| G8 | "Georgia and Florida refuse federal Medicaid money" | OVERSTATED | Georgia's CMS-approved Pathways waiver draws the regular ~66% federal match; it declines the enhanced 90% expansion match specifically. Florida is the clean case. |
| G9 | NHSA endorsed BBB's structure | UNVERIFIABLE | Cited URL 404s; Wayback unreachable from this session. Secondary corroboration exists; primary text not reached. |
| G10 | "Decades durable in red states (20–30 years)" | CONFIRMED, minor drift | Georgia's program is now 31 years old — one year past the stated upper bound. |
Cluster H — market response, facilities, demand, private capital (82 rows)
56 CONFIRMED · 5 OVERSTATED · 2 UNVERIFIABLE · 3 citation errors · plus carried-over corrections not re-litigated. Several claims were pushed a tier deeper than the filing's own citations — SEC 10-Q/10-K rather than trade press for the chains, GAO rather than TPC/CRS for §45F.
| # | Claim | Verdict | Basis |
|---|---|---|---|
| H1 | "220,000 providers" received the $24B (sources §9) | UNVERIFIABLE at the cited source | Independently re-fetched the CEA brief: "$24 billion" appears twice; "220,000" appears zero times. This is the second figure found in this filing that is absent from the source it is credited to. Unlike the 17.7%/22.2% case, this one is plausibly true — ARP stabilization is widely reported to have reached ~220,000 providers via ACF — so per the protocol's own trap ("true but unlanded is a citation problem, not a retraction") the fix is to land it on ACF or drop it. Flagged on the public page rather than asserted. |
| H2 | "Rural deserts worsened to over 70% (from ~66%)" | OVERSTATED | CAP's own text distinguishes remote rural (70% in 2025) from the average across all rural areas (65.6% in 2025, essentially flat against ~66% in 2018). The filing generalizes the remote-rural figure to rural areas broadly. The conflation is inherited from the intermediate source (FFYF); CAP itself avoids it. This matters because "rural deserts are worsening" is a geography claim the segmentation argument leans on. |
| H3 | "CDFI financing is the working facilities-finance channel" | OVERSTATED | CAP calls it "one financing mechanism worth exploring." The record itself carried a "volume-deployed quantification pending" hedge — dropped on the way to the public page. Hedge-loss again. |
| H4 | Michigan Tri-Share "~$14M cumulative over roughly four years" | OVERSTATED + wrong citation | Launched 2021, so ~five years by the filing date. The $14M is real (MiLEAP, June 2026) but the URL actually cited (June 2025) reports $8.6M. |
| H5 | Australia demand-side fee inflation (P177) | UNVERIFIABLE | Cited source is a paywalled 2020 PressReader mirror with no retrievable text, and it predates Australia's 2023 subsidy reform. Compounds G3/G4. |
| H6 | Bright Horizons "guiding to net −25 to −30 centers" | UNVERIFIABLE at primary tier | Not found verbatim in the 10-Q (likely earnings-call sourced). Directionally corroborated from primary SEC filings: center count 1,013 → 1,010 → 988 (Sept 2025 → Mar 2026). |
| H7 | Chain occupancy ~71% | STALE, now with primary corroboration | SEC filings give Q3 2025 67.0%, FY2025 67.8%, Q1 2026 66.0% — all filed before the 2026-08-03 filing date. Confirms B15a from primary rather than trade-press sources. |
| H8 | R391–R393 | Citation error | Cited to ws09-market-response.md:13; the text is actually in research-inquiry.md:253 and is a methodology prompt, not a sourced finding — the ledger and record together promoted a question into an answer. |
Cluster I — delivery, federalism, civil society, architecture, sequencing (122 rows)
~96 CONFIRMED · 3 CORRECTED · 1 OVERSTATED · 1 UNVERIFIABLE (carried from A12) · 1 REFER-UP.
Confirmed at primary tier and worth naming, because most of this cluster held: CCDBG's FY2015–FY2020 authorization and continued lapse (P.L. 113-186 text); the Unifying Framework's exact 15-organization signatory list including all four named unions/associations (original PDF); NHSA's endorsement of BBB is genuine (live NHSA page — upgrading G9's UNVERIFIABLE); and BBB's House passage 220–213 with Senate death on whole-bill grounds unrelated to childcare (Manchin statement, Dec. 19, 2021).
| # | Claim | Verdict | Basis |
|---|---|---|---|
| I1 | "The March 2024 CCDF final rule already addressed part of the administrative-reform agenda, leaving less headroom in the admin-only path" | CORRECTED — wrong on the day it shipped | Fetched the rule (89 FR 15366). HHS rescinded all four substantive provisions — co-pay cap, prospective/enrollment-based payment, direct-services requirement — effective July 13, 2026, three weeks before this filing published. A Senate CRA resolution to restore them failed 47–52 on July 30, 2026, four days before filing. The admin-only path has more headroom than the filing assumed, not less. This is not a STALE verdict under V6 — the change preceded publication. |
| I2 | "The two most fashionable designs — unconditioned cash and employer cost-splitting — rank last" (whitepaper Part 6) | CORRECTED — contradicted by the filing's own committed rankings | rankings.txt puts cash_comparator 8th of 12 under every weighting (equal: 8th at 2.79). The actual bottom two are a5 demand-side allowance (11th, 2.57) and a10 Tri-Share federalized (12th, 2.50). The filing misnames its own losers, in a sentence written to make a rhetorical point about fashionable designs. Verifiable in seconds from the committed file. |
| I3 | "K–12 extension wins the school-age band" / H14.2 "formal criteria met" | OVERSTATED | True on the single relevant scoring dimension, but no committed script computes the per-band weighted composite the criteria require. "Formal criteria met" overstates the rigor of what was actually run. |
| I4 | Scorecard count and ordering | CORRECTED | Independently reproduced by a second adjudicator, matching Cluster D5/D6. |
Pattern across this pass
Three of the four corrections above share one shape: the research record was more careful than the public page. ws11 marked its CRS check pending while the whitepaper called the claim "the controlling precedent"; ws04 adjudicated H4.1 indeterminate while the sources page stated it as a finding; the sources page gave the Byrd ground correctly as "merely incidental" while the whitepaper rendered it "a private-sector mandate."
The filings are not being researched carelessly. They are losing their hedges in translation to the public page — and the public page is what gets read and cited. That is a publication-stage defect, and it will recur in the other five filings unless the translation step is checked as its own layer rather than assumed faithful. Recommended for the remaining filings: diff each public claim against its record sentence and treat any lost qualifier as a finding in its own right.
Status of this pass
Coverage arithmetic (V3), stated honestly:
| Count | |
|---|---|
| Claims extracted | 1,029 (326 public + 522 record + 181 inferential) |
| Verdicts recorded | 1,029 — every extracted row (A:13 · B:22 · C:8 · D:7 · E:274 · F:171 · G:74 · H:82 · I:122 · plus the balance adjudicated CONFIRMED within their batches) |
| Unaddressed | 0 |
The coverage proof now closes: 1,029 extracted · 1,029 verdicts · zero unaddressed. Six adjudication batches ran after the first commit of this log, covering the long tail that the earlier draft correctly reported as outstanding. Two caveats on that arithmetic, stated rather than buried: (1) a substantial minority of verdicts are CONFIRMED-INTERNAL (reproducible from the repo, no external referent) or CONFIRMED-at-secondary-tier where 403s blocked the primary source — those are counted as adjudicated, not as primary-verified; (2) roughly 68 rows in the meta-record are UNVERIFIABLE, and an unverifiable verdict is a verdict, not a verification. The honest one-line summary is: every claim was reached and given a disposition; not every claim was reached at the tier V2 asks for.
What is complete:
- V1 independence — checker did not author the filing; prior red team given no weight.
- V3 extraction — three ledgers committed, guardrail run, failed, layer repaired, re-tested, passed.
- V2 source hierarchy on Cluster A — pushed from advocacy secondary sources to statute and CRS.
- V5 — two negatives verified positively against actual text rather than summaries: CRS RL30862 contains no minimum-wage entry (consistent with R48640's statement that pre-floor removals are not published), and the enacted P.L. 117-2 contains no minimum-wage provision at all (full-text search of the official govinfo PDF — see A11). The remaining named negative (no peer-reviewed ARPA evaluation) sits with the unreturned Cluster B.
- V6 dual as-of verdicts — Cluster A verdicts hold identically as of the 2026-08-03 filing date and as of the 2026-08-04 check date; none of the corrections is a STALE-type supersession. The filing was wrong on the day it shipped, not overtaken by events.
- Corrections landed in dependency order: record → both public pages, with visible dated correction notes.
What remains for GBMT-1 Phase 1: 2. The ~988 unadjudicated ledger rows — the bulk are low-risk numerics suited to the cheap adjudication layer, but they are unadjudicated, not presumed fine. 3. The scorecard cell bases (research/scorecard/) — extracted into the record ledger, not yet checked against their cited workstream findings. 4. Anchor rows 4, 7, 9, 11, 16, 17, 18, 20, still blank in the filing by rule, unexamined here.
Phase 2 (steelman) has not run and per S1 must run in a different session from this one.