Protocol: method/verification-protocol.md, Phase 2 (S1–S5). Filing under test: Whitepaper No. 1, filed 2026-08-03, as corrected by Phase 1. Date: 2026-08-04.
Deviation from S1, logged not hidden. S1 requires Phase 2 to run in a session different from the fact-check, so the steelman is not anchored by the checker's conclusions. This pass was run in the same session at the user's instruction. Mitigation: every steelman was constructed in a fresh agent context that received the filing and the corrected record but not the checker's reasoning about where the filing was weak, and each was required to derive the tilt itself per S2. That preserves S1's purpose — independence of the builder from the checker's conclusions — while breaking its letter. A reader who thinks that insufficient should discount this phase accordingly.
Second deviation: S1 also requires Phase 1's corrections to have landed first. The major corrections had landed (the Byrd cluster, the numeric pass); Phase 1's long-tail adjudication was still in flight when Phase 2 was launched. Any steelman premise later undermined by a late Phase 1 verdict is flagged in the adjudication below.
S2 — The tilt, derived by the adjudicator from the filing's own text
Written before reading any steelman output, so it can be compared against three independently-derived tilt statements.
Which direction the filing leans. Its conclusion is that universal childcare is feasible as a design and sequencing problem: "pay the workforce first, build supply second, extend benefits third." The binding constraint is labour, and the lever is compensation.
Which architecture won. All four top-ranked designs — Head Start scaled to universal, supply-first build-out, public option, federal fallback ladder — are supply-side, federal, and direct-provision. The two designs ranked last are the two demand-side/private-channel instruments: unconditioned cash and employer cost-splitting.
Whose framing it adopted. The provider-and-workforce frame: care as a service to be publicly provisioned, whose central problem is the compensation of the people delivering it. Families appear mostly as demand to be met, rarely as agents choosing among arrangements.
What it argued against, scored down, or never entertained. Three things, and the third is the most consequential:
- Instrument choice. Cash and employer channels are ranked last, and the filing's demand-side cautionary tale is a single case (Australia). Whether the scorecard's dimensions can even permit a cash instrument to win is not examined.
- The diagnosis itself. "Workforce binds" is asserted against a model whose output is a direct function of assumed staff-child ratios. The competing diagnosis — that ratios and credentialing are the binding constraint — is not scored as an architecture at all.
- Whether the goal is right. The filing asks only whether universal coverage is achievable. No hypothesis in
adjudication-criteria.mdis capable of returning "this should not be done," and Quebec — the longest-running universal program available — is cited only for its waitlist, not for its child-outcome literature.
The revealing detail (S2's instruction to derive the tilt from the text, not the subject). The filing did run a steelman: steelman-feasibility.md. But it steelmanned feasibility — the proposition that this can be done — against its own uniformly skeptical hypothesis set. That tells you what the filing treated as the unfashionable direction: pessimism about delivery. It never steelmanned a different instrument or a different goal. Its own red team, Attack 5, concedes the point in the filing's own words: "All four winners embody the same analyst's pass-1 conclusions (supply-side, federal, direct provision); if the scores encode the priors, stability across weightings measures internal consistency, not truth."
That concession is the strongest internal warrant for this phase. The filing identified the risk and then published the scorecard anyway, with an independent re-score listed as a standing requirement.
Therefore the three steelman targets, assigned to independent builders:
| # | Target | Why it is the unfashionable direction here |
|---|---|---|
| A | Cash / employer instruments beat supply-side provision | The filing ranks them last and never tests whether its scales permit them to win |
| B | The binding constraint is regulatory (ratios, credentials), not compensation | Not scored as an architecture; the headline is a function of the model's ratio parameters |
| C | Universal non-parental care for 0–2 is not clearly beneficial | The filing's framing cannot return this answer; Quebec's outcome literature is unexamined |
S3 / S4 — Steelman A: cash and employer instruments
Independently-derived tilt (builder's own words), which matches the adjudicator's: the filing "leans toward publicly financed supply of formal, center-anchored, credentialed care with money entering on the provider side," adopting "the ECE professional field's framing: care measured in credentialed FTEs, quality operationalised as inputs… and never as measured child outcomes."
A-1. The scorecard's dimensions encode the answer — STEELMAN WINS
Three of the fourteen dimensions do not score an outcome. They score the instrument category, using the architecture's own name as the scale anchor. Verbatim from scales.md:
| Dimension | 1 means | 5 means |
|---|---|---|
relief_supply |
"Builds no capacity; pure transfer" | "Builds new capacity" |
passthrough_risk |
"Pure demand-side into inelastic supply" | "Supply-side operating grants / direct provision" |
integrity_risk |
"Netherlands-shaped: mass demand-side clawback exposure" | "Direct provision / small-N grantee audit surface" |
A cash or voucher design is scored at or near 1 on all three by definition, before any evidence is consulted. "Demand-side" is not a finding about the instrument; it is the printed anchor text. The filing then reports that supply-side designs win and that "the two most fashionable designs — unconditioned cash and employer cost-splitting — rank last."
Verified by the adjudicator, from the repo alone. Dropping those three dimensions and re-running the board — reproducible with rank-sensitivity.py, which reproduces the committed rankings before computing any variant:
| Weighting | Committed top 4 | Top 4 without instrument-labelled dimensions |
|---|---|---|
| equal | a8 · a3 · a4 · a6 | a8 · a3 (exact tie, 3.73) · a4 · a11 caregiver-choice — a6 falls to 7th |
| labor_supply_first | a8 · a4 · a3 · a6 | a8 · a4 · a3 · a11 — a6 falls to 6th |
| child_dev_first | a8 · a3 · a4 · a6 | a8 · a3 · a4 · a11 — a6 falls to 5th |
| equity_first | a8 · a3 · a4 · a11 | a11 · a3 · a8 · cash comparator — a4 falls to 6th |
Corrected 2026-08-11. As first written this table ran against the pre-2026-08-07 board and was never recomputed after the independent blind re-score moved 47 cells (scorecard/rescore-log.md) — it was committed 2026-08-08 and published 2026-08-10, both after that re-score. The stale version showed a3 first in both columns under all four weightings. The structural finding survives the recomputation and is sharpened by it; the claims made about a3 do not, and are corrected throughout this section.
a6 supply-first — the flipping fourth seat — leaves the top four under every weighting, and a11 caregiver-choice, a home-based demand-side design, takes the vacated seat in three of the four. Under equity-first the effect is sharper still: a11 leads outright, and the cash comparator — a design the anchor text cannot score above the floor — enters the top four at 3.50 while a4 drops to 6th. The all-weightings stable set moves from {a3, a4, a8} to {a3, a8, a11}. a3 leads nothing in either specification: a8 leads all four weightings on the committed board and three of four here, tying a3 exactly under equal.
Verdict: the steelman wins this point outright, and it is verifiable without leaving the repository. Per S4 this is applied to the record, not footnoted: the "four designs survive every ranking" claim is conditional on three dimensions that cannot register a demand-side design as anything but a failure.
A-2. The child-outcome omission — STEELMAN PARTIALLY WINS; its stated version overreached
The builder wrote that the filing "cites none of" HSIS, Tennessee, Boston, or the Quebec outcome literature. That is wrong, and the adjudicator corrects it: ws07-quality.md engages Tennessee VPK ("the negative case") and Boston ("the positive case") substantively, and both appear on the public pages. A steelman is not licensed to overstate either.
What survives is sharper and worse:
- The Head Start Impact Study appears nowhere in the record — zero hits across
childcare/research/andchildcare/docs/. The filing's own inquiry protocol commissioned it by name: §210 asks for "Evidence on outcomes from large-scale programs: Head Start (including the Impact Study and the fade-out debate)." The filing ranks scaled Head Start in the top four under every weighting — second on the committed board under three of the four — and never engages the largest randomized evaluation of Head Start. - Baker–Gruber–Milligan appears only as a citation count.
phase0-findings.md:38records the OpenAlex search returning "Baker–Gruber–Milligan (NBER w11832; 271 citations)" — the literature search found the canonical Quebec outcome paper and logged its popularity, never its finding. The protocol §210 also asked for it by name, including "subsequent work contests scope and duration — represent the disagreement accurately." - The protocol likewise commissioned QRIS evaluation evidence and "the regulation–supply tradeoff." Neither was delivered.
Verdict: OVERSTATED as the builder framed it; UPHELD in the narrower form. The filing scores a child_dev_first weighting across twelve architectures without a child-outcome dimension, and its own protocol commissioned outcome work that the research pass did not perform. That is a scope failure the filing never disclosed.
A-3. Australia as the demand-side cautionary tale — STEELMAN WINS, already landed
Independently reached by the precedents adjudication in Phase 1 (Cluster G3/G4): the ACCC's own December 2023 report found the rate cap had "only limited effectiveness," and "+4.4% in a representative year" is an unsourced gloss on a single quarter. Both corrected on the public pages. That the steelman and the fact-check converged on this from different directions raises confidence in it.
A-4. The employer channel cleared the filing's own procedural test — NOT ADJUDICATED
The builder reports 26 U.S.C. § 45F as amended by Pub. L. 119-21 § 70401 (July 4, 2025) — a reconciliation act — adding an "intermediate entity" contract path matching Tri-Share's mechanism. If so, the filing's a10 procedural = 3 understates a channel that has already passed through reconciliation. Left unadjudicated: the adjudicator did not independently reach the statute this pass. Recorded as an open item, not a verdict.
A-5. What defeats the steelman, in its own evidence
Stated because S4 requires recording what a surviving conclusion beat. Baby's First Years (Gennetian et al., Nature Human Behaviour 8:1514–1529) — an RCT of unconditional cash to mothers of infants — found no statistically significant difference in mothers' paid work or children's time in childcare. Cash at plausible magnitudes has a measured null on precisely the objective the filing's labor_supply_first weighting scores. The builder found this itself and reported it against its own case, which is what an honest steelman looks like.
Overall verdict on Steelman A: partially survives (S4 outcome 2). The filing's recommendation to fund supply is not overturned by this point — a3 and a8 hold top-four places under every weighting in both specifications. What falls is the claim that four designs survive every ranking; the claim that cash and employer instruments rank last on the evidence (they rank last partly on definitional grounds); and any claim that a3 leads, which it does under no weighting in either specification — a8 leads all four on the committed board.
S5 — Scorecard consequences
Two independent sensitivities now exist, both published in scorecard/rationale.md and above:
- Under Phase 1's corrected Byrd finding (procedural re-score, computed on the pre-2026-08-07 board): top-4 set unchanged; a8 ties a3 under
child_dev_firstand leads underlabor_supply_first. Superseded as a statement about a3's primacy — the blind re-score has since put a8 first under all four weightings outright, so there is no a3 lead left for this sensitivity to weaken. - Dropping instrument-labelled dimensions: top-4 set changes in three of four weightings — a6 out, a11 in.
Does the recommendation move? Partly. "Fund supply, fund compensation, house standards inside federal spending programs" survives both. "These four architectures, stable under every weighting" does not survive the second. The defensible published claim is: a8 (public option) leads under every weighting on the committed board; a3 (Head Start scaled) holds a top-four place under every weighting in both specifications but leads none of them; and the remainder of the top tier is weighting- and specification-dependent, with one member (a6) an artifact of three dimensions that score instrument type rather than outcome.
S3 / S4 — Steelman C: is universal 0–2 non-parental care desirable at all?
Independently-derived tilt: "pro-universal-coverage on the goal, skeptical on the means. Universal coverage of all 51.2M children 0–12 enters as the definition of the problem." The builder reached the same observation the adjudicator did about the filing's own steelman — it argues the build is more achievable, so "the unfashionable side it identified was optimism about the build, never pessimism about the goal."
C-1. The apparatus cannot express the question — STEELMAN WINS, verified from the repo
Three findings, all checkable without leaving the repository, all verified by the adjudicator:
- No dimension measures child development. All fourteen dimensions in
scales.mdscore money, coverage, procedure, federalism, durability, pass-through, participation, preference fit, integrity, distribution, evaluability. None scores an effect on children. - The
child_dev_firstobjective is implemented as coverage.rank.py, verbatim:"child_dev_first": {…, "relief_workforce": 3, "coverage_34": 2, "preference_fit": 2, "evaluability": 2, "coverage_02": 2}. The "child development" weighting scores workforce relief highest (×3) and otherwise rewards covering more children. A design scores better for child development by enrolling more under-2s — which is precisely the contested proposition, entered as an axiom of the objective function rather than tested. - No hypothesis could return "don't do this."
adjudication-criteria.mdhas no supported-branch meaning the goal is wrong.
C-2. The filing's own protocol review demanded this work, named it, and it was not done — STEELMAN WINS
This is the most damaging item in Phase 2, and it is entirely internal. protocol-review.md:
- Line 36: "The protocol never states what the program optimizes… Quebec optimized labor supply and is dinged on child outcomes."
- Line 114, recommendation 8: "Adversarial presentation for polarized literatures. Where the field genuinely disagrees (Quebec child outcomes, Tennessee VPK, Head Start fade-out), present each side's preferred specification and what evidence would resolve the dispute — never a settled-sounding sentence."
Three literatures named. The filing delivered one. ws07-quality.md presents Tennessee VPK adversarially against Boston — exactly as instructed. Quebec child outcomes and Head Start fade-out were not delivered at all:
phase0-findings.md:38records the OpenAlex search returning Baker–Gruber–Milligan, the BGM follow-up, and Cornelissen et al. "on the first query." The papers were found on day one and logged as a citation count. Quebec appears in the filing only as a waitlist statistic.- The Head Start Impact Study appears nowhere in the record (zero hits), while scaled Head Start ranks top-four under every weighting.
The protocol review predicted the omission, named the exact three literatures, the search retrieved the papers immediately, and two of the three were dropped. That is not a sourcing gap; it is a specified deliverable that was not performed and not disclosed — nothing in the deviations log records it.
C-3. The substantive case — PARTIALLY SURVIVES, and the builder said so first
Marshalled at the required tier: BGM (JPE 2008; AEJ:Policy 2019) on persistent non-cognitive harms and crime effects; Havnes–Mogstad (JPubE 2015) showing the same Norwegian reform is positive to the 69th percentile and significantly negative above the 81st; Cornelissen et al. (JPE 2018) finding reverse selection on gains; Fort–Ichino–Zanella (JPE 2020) finding IQ costs at 0–2 rising with family income in a 1:4/1:6-ratio system, which forecloses the "just low quality" defence; Drange–Havnes (JOLE 2019), a real randomized lottery, finding no effect among high-income families.
The builder then reported the evidence that defeats its own strong claim, which is what S3 asks for: Montpetit, Carrer & Beauregard (April 2026) find BGM's behavioural harms "do not translate into lower educational attainment or reduced earnings" and estimate MVPF 2.83; Kottelenberg–Lehrer (JOLE 2017) find gains for disadvantaged single-parent children, inverting the gradient's location; Herbst (JOLE 2017) finds the US Lanham Act — universal — produced +0.36 SD concentrated among the most disadvantaged.
Verdict: partially survives (S4 outcome 2). It does not establish that universal 0–2 care is harmful, and it does not claim to. It establishes that the question is live, that the field is genuinely divided, and that the filing's instrument cannot represent either side.
C-4. The result that matters most — a right answer resting on a wrong reason
The steelman vindicates the filing's 0–2 architecture while dissolving its justification. ws14/ws16 reach "supply-first plus caregiver-choice blend, not centre-led" for 0–2 via revealed preference and supply economics. The child-outcome literature arrives at the same destination by an entirely different route: returns are highest where the home counterfactual is weakest, and kin/home settings preserve near-1:1 ratios. The filing is right about 0–2 for reasons it never examined — which makes the conclusion fragile to precisely the objection it never heard.
C-5. S5 — and an honest refusal
Verified by the adjudicator: setting coverage_02 to weight 0 in child_dev_first — merely declining to assume that covering under-2s benefits children — leaves the top-4 set unchanged (a3 3.89 · a8 3.83 · a6 3.83 · a9 3.56). The recommendation survives agnosticism.
The builder computed that inverting the dimension changes the set (a9 out, a4 in) and then declined to offer it as the correct re-score, because inversion also demotes a11 caregiver-choice — the architecture its own evidence most supports. That refusal is the right call and it sharpens the finding: the scorecard has no dimension capable of carrying this question, and forcing it into coverage_02 corrupts a dimension that means something else. The fix is a new child_effect_0_2 dimension, not a re-weighting. Recorded as a recommendation, not applied — adding dimensions to a published scorecard is the filing's call, not the checker's.
Convergence across independent builds
Steelmen A and C were built in separate contexts, on different targets, with no access to each other. Both independently arrived at the same structural defect from opposite ends: the scorecard's dimensions encode the filing's conclusions. A found three dimensions whose scale anchors name the instrument category; C found that the objective meant to represent child development contains no child-outcome term. Both also independently surfaced the same omission — the Head Start Impact Study — which the filing's own protocol review had named.
Independent convergence on a defect neither was pointed at is the strongest evidence this phase produced.
S3 / S4 — Steelman B: is the constraint regulatory rather than a factor price?
Independently-derived tilt: the filing holds "a factor-price diagnosis of scarcity" — care is scarce because providers are underpaid, so the fix is money routed to compensation. The builder then showed the alternative was excluded by construction, not argued against:
deregulat*,relax*,loosen*return zero hits repo-wide.ratioappears only as a modelling input, never as a policy variable.- None of the twelve scored architectures alters the regulatory production function.
- No scorecard dimension measures cost-per-slot or regulatory burden — so even a deregulatory design could not have been scored.
research-inquiry.md§161–162 commissioned work on "square footage per child, outdoor space, egress and fire code for infants… zoning and land use as an obstacle."ws05-facilities.mdis twelve lines on deserts and CDFI financing and mentions no facility code. A third commissioned-and-dropped workstream, alongside the two in C-2.
B-1. The headline number is a function of a parameter the filing never examined — STEELMAN WINS
fte_model.py treats staff-child ratios as an exogenous technological constant. Verified by the adjudicator, re-running the committed model with only the ratio parameters changed and every other value at the filing's own settings:
| Ratio regime | FTEs required | Net new | Hires/yr |
|---|---|---|---|
| Filing central (published) — r = (4.5, 10.0, 14.0) | 2,807k | 1,757k | 798k |
| Florida statute (Fla. Stat. § 402.305(4)) — r ≈ (7, 17, 25) | 1,727k | 677k | 501k |
| Looser still (Mississippi-shaped) | 1,623k | 573k | 473k |
The decisive comparison. Halving turnover — the filing's own "arithmetically dominant lever," the thing Part 2 is built around — takes hiring from 798k/yr to 509k/yr. Adopting the ratios Florida's own statute already permits, with no pay rise at all, takes it to 501k/yr. The two levers are equivalent in magnitude; one costs $45.9B/yr and the other costs nothing and needs no federal vehicle.
A second point sharpens it: the model's r02 = 4.5 is tighter than the median state's law and tighter than NAEYC's own recommendation, while the docstring describes the values as "licensing-typical." And the filing praises Florida in Part 5 as a durable multi-decade success while modelling ratios roughly 1.7–1.8× tighter than Florida requires.
B-2. "Binding" fails on the filing's own definition — STEELMAN WINS
research-inquiry.md defines the term: a constraint is binding if relaxing it alone would materially raise achievable output and relaxing the others would not. Relaxing the ratio alone does approximately what relaxing compensation does. The second clause fails on the filing's own test. The word doing the most work in the whitepaper — stamped Binding constraint in Part 2 — is not established.
B-3. What the regulation buys — STEELMAN PARTLY WINS
Marshalled at tier: Hotz & Xiao (AER 101(5) 2011) — one fewer infant per staffer reduces centres in the average market 9.2–10.8%, concentrated in poorer markets; Ali, Herbst & Makridis (IZA DP 14684), the best-identified design either side has — new group-size regulation cuts childcare job postings 5.5–8.4% and lowers maternal LFP 2.0pp, which inverts the filing's model by making tighter ratios reduce childcare employment; Perlman et al. (PLOS ONE 2017) — within permissible ranges, ratio variation shows pooled r = 0.03 with child outcomes, threshold story tested and rejected; Vanderbilt's Prenatal-to-3 Center rates ratios "Needs Further Study." And the filing's own §7 supplies the point: what separated Boston from Tennessee was "curriculum + coaching + pay near parity" — all process quality, not one structural ratio.
B-4. Where the steelman breaks, in its own words
Recorded because S4 requires it, and because the builder volunteered all three:
- Currie & Hotz (JHE 2004): caregiver training beyond high school lowers non-vehicular accidental child deaths ~18% with clean placebo behaviour — the best causal estimate in the file, and it defends credentials. The same paper finds looser ratios in family homes raise accident rates — which is exactly the 0–2 modal setting the filing recommends building on.
- The builder's own cross-state test found looser-ratio states have fewer ECE workers per child (r = −0.49), with a confound identified but a confidence interval hugging the boundary.
- The childcare wage relative to its own state's median is uncorrelated with ratio stringency (r = −0.089) — so the steelman cannot claim regulation caused the low pay. The filing's structural reading of the wage wins that sub-argument outright.
Verdict on Steelman B: partially survives (S4 outcome 2) — on the load-bearing word
The filing is right that at its assumed ratios the required flow is unprecedented; right about the five-country pattern; right that low pay drives turnover. What does not survive is "binding" as a claim about the world rather than about a parameter. Applied to the record per S4, not footnoted.
S5 consequence
The scorecard is structurally incapable of testing this: it contains no regulatory-reform architecture and no cost-per-slot or regulatory-burden dimension. No re-score can be offered, because there is nothing to re-score — which is itself the finding. Recorded as a recommendation: a future pass needs a regulatory_reform architecture and a cost_per_slot dimension before this question can be scored at all.
Convergence across all three builds
Three cases, three targets, no access to each other. All three land on the same structural fact from different directions: the filing's instrument encodes its conclusions. A found three dimensions anchored on instrument type; B found no dimension capable of registering regulatory cost and no architecture that alters it; C found the child-development objective built entirely from coverage terms. And all three found commissioned work that was never delivered and never logged — Quebec child outcomes and Head Start fade-out (C), facilities codes and zoning (B), QRIS evaluation evidence (both).
The pattern, stated once: GBMT-1's protocol was better than its execution, and the gap between them is not recorded anywhere in the filing. That is the finding this phase exists to produce, and it is more useful than any single corrected number.