GUBMENTPlain talk · policy frontier
Filings / Childcare / Sources / GBMT-1 (Childcare) — Phase 2 Steelman
GBMT-1 · Research record · No. 1

GBMT-1 (Childcare) — Phase 2 Steelman Log

childcare/research/steelman-log.md
This is a working research document from the childcare filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 1, not a summary of it.

Protocol: method/verification-protocol.md, Phase 2 (S1–S5). Filing under test: Whitepaper No. 1, filed 2026-08-03, as corrected by Phase 1. Date: 2026-08-04.

Deviation from S1, logged not hidden. S1 requires Phase 2 to run in a session different from the fact-check, so the steelman is not anchored by the checker's conclusions. This pass was run in the same session at the user's instruction. Mitigation: every steelman was constructed in a fresh agent context that received the filing and the corrected record but not the checker's reasoning about where the filing was weak, and each was required to derive the tilt itself per S2. That preserves S1's purpose — independence of the builder from the checker's conclusions — while breaking its letter. A reader who thinks that insufficient should discount this phase accordingly.

Second deviation: S1 also requires Phase 1's corrections to have landed first. The major corrections had landed (the Byrd cluster, the numeric pass); Phase 1's long-tail adjudication was still in flight when Phase 2 was launched. Any steelman premise later undermined by a late Phase 1 verdict is flagged in the adjudication below.


S2 — The tilt, derived by the adjudicator from the filing's own text

Written before reading any steelman output, so it can be compared against three independently-derived tilt statements.

Which direction the filing leans. Its conclusion is that universal childcare is feasible as a design and sequencing problem: "pay the workforce first, build supply second, extend benefits third." The binding constraint is labour, and the lever is compensation.

Which architecture won. All four top-ranked designs — Head Start scaled to universal, supply-first build-out, public option, federal fallback ladder — are supply-side, federal, and direct-provision. The two designs ranked last are the two demand-side/private-channel instruments: unconditioned cash and employer cost-splitting.

Whose framing it adopted. The provider-and-workforce frame: care as a service to be publicly provisioned, whose central problem is the compensation of the people delivering it. Families appear mostly as demand to be met, rarely as agents choosing among arrangements.

What it argued against, scored down, or never entertained. Three things, and the third is the most consequential:

  1. Instrument choice. Cash and employer channels are ranked last, and the filing's demand-side cautionary tale is a single case (Australia). Whether the scorecard's dimensions can even permit a cash instrument to win is not examined.
  2. The diagnosis itself. "Workforce binds" is asserted against a model whose output is a direct function of assumed staff-child ratios. The competing diagnosis — that ratios and credentialing are the binding constraint — is not scored as an architecture at all.
  3. Whether the goal is right. The filing asks only whether universal coverage is achievable. No hypothesis in adjudication-criteria.md is capable of returning "this should not be done," and Quebec — the longest-running universal program available — is cited only for its waitlist, not for its child-outcome literature.

The revealing detail (S2's instruction to derive the tilt from the text, not the subject). The filing did run a steelman: steelman-feasibility.md. But it steelmanned feasibility — the proposition that this can be done — against its own uniformly skeptical hypothesis set. That tells you what the filing treated as the unfashionable direction: pessimism about delivery. It never steelmanned a different instrument or a different goal. Its own red team, Attack 5, concedes the point in the filing's own words: "All four winners embody the same analyst's pass-1 conclusions (supply-side, federal, direct provision); if the scores encode the priors, stability across weightings measures internal consistency, not truth."

That concession is the strongest internal warrant for this phase. The filing identified the risk and then published the scorecard anyway, with an independent re-score listed as a standing requirement.

Therefore the three steelman targets, assigned to independent builders:

# Target Why it is the unfashionable direction here
A Cash / employer instruments beat supply-side provision The filing ranks them last and never tests whether its scales permit them to win
B The binding constraint is regulatory (ratios, credentials), not compensation Not scored as an architecture; the headline is a function of the model's ratio parameters
C Universal non-parental care for 0–2 is not clearly beneficial The filing's framing cannot return this answer; Quebec's outcome literature is unexamined

S3 / S4 — Steelman A: cash and employer instruments

Independently-derived tilt (builder's own words), which matches the adjudicator's: the filing "leans toward publicly financed supply of formal, center-anchored, credentialed care with money entering on the provider side," adopting "the ECE professional field's framing: care measured in credentialed FTEs, quality operationalised as inputs… and never as measured child outcomes."

A-1. The scorecard's dimensions encode the answer — STEELMAN WINS

Three of the fourteen dimensions do not score an outcome. They score the instrument category, using the architecture's own name as the scale anchor. Verbatim from scales.md:

Dimension 1 means 5 means
relief_supply "Builds no capacity; pure transfer" "Builds new capacity"
passthrough_risk "Pure demand-side into inelastic supply" "Supply-side operating grants / direct provision"
integrity_risk "Netherlands-shaped: mass demand-side clawback exposure" "Direct provision / small-N grantee audit surface"

A cash or voucher design is scored at or near 1 on all three by definition, before any evidence is consulted. "Demand-side" is not a finding about the instrument; it is the printed anchor text. The filing then reports that supply-side designs win and that "the two most fashionable designs — unconditioned cash and employer cost-splitting — rank last."

Verified by the adjudicator, from the repo alone. Dropping those three dimensions and re-running the board — reproducible with rank-sensitivity.py, which reproduces the committed rankings before computing any variant:

Weighting Committed top 4 Top 4 without instrument-labelled dimensions
equal a8 · a3 · a4 · a6 a8 · a3 (exact tie, 3.73) · a4 · a11 caregiver-choicea6 falls to 7th
labor_supply_first a8 · a4 · a3 · a6 a8 · a4 · a3 · a11a6 falls to 6th
child_dev_first a8 · a3 · a4 · a6 a8 · a3 · a4 · a11a6 falls to 5th
equity_first a8 · a3 · a4 · a11 a11 · a3 · a8 · cash comparatora4 falls to 6th

Corrected 2026-08-11. As first written this table ran against the pre-2026-08-07 board and was never recomputed after the independent blind re-score moved 47 cells (scorecard/rescore-log.md) — it was committed 2026-08-08 and published 2026-08-10, both after that re-score. The stale version showed a3 first in both columns under all four weightings. The structural finding survives the recomputation and is sharpened by it; the claims made about a3 do not, and are corrected throughout this section.

a6 supply-first — the flipping fourth seat — leaves the top four under every weighting, and a11 caregiver-choice, a home-based demand-side design, takes the vacated seat in three of the four. Under equity-first the effect is sharper still: a11 leads outright, and the cash comparator — a design the anchor text cannot score above the floor — enters the top four at 3.50 while a4 drops to 6th. The all-weightings stable set moves from {a3, a4, a8} to {a3, a8, a11}. a3 leads nothing in either specification: a8 leads all four weightings on the committed board and three of four here, tying a3 exactly under equal.

Verdict: the steelman wins this point outright, and it is verifiable without leaving the repository. Per S4 this is applied to the record, not footnoted: the "four designs survive every ranking" claim is conditional on three dimensions that cannot register a demand-side design as anything but a failure.

A-2. The child-outcome omission — STEELMAN PARTIALLY WINS; its stated version overreached

The builder wrote that the filing "cites none of" HSIS, Tennessee, Boston, or the Quebec outcome literature. That is wrong, and the adjudicator corrects it: ws07-quality.md engages Tennessee VPK ("the negative case") and Boston ("the positive case") substantively, and both appear on the public pages. A steelman is not licensed to overstate either.

What survives is sharper and worse:

Verdict: OVERSTATED as the builder framed it; UPHELD in the narrower form. The filing scores a child_dev_first weighting across twelve architectures without a child-outcome dimension, and its own protocol commissioned outcome work that the research pass did not perform. That is a scope failure the filing never disclosed.

A-3. Australia as the demand-side cautionary tale — STEELMAN WINS, already landed

Independently reached by the precedents adjudication in Phase 1 (Cluster G3/G4): the ACCC's own December 2023 report found the rate cap had "only limited effectiveness," and "+4.4% in a representative year" is an unsourced gloss on a single quarter. Both corrected on the public pages. That the steelman and the fact-check converged on this from different directions raises confidence in it.

A-4. The employer channel cleared the filing's own procedural test — NOT ADJUDICATED

The builder reports 26 U.S.C. § 45F as amended by Pub. L. 119-21 § 70401 (July 4, 2025) — a reconciliation act — adding an "intermediate entity" contract path matching Tri-Share's mechanism. If so, the filing's a10 procedural = 3 understates a channel that has already passed through reconciliation. Left unadjudicated: the adjudicator did not independently reach the statute this pass. Recorded as an open item, not a verdict.

A-5. What defeats the steelman, in its own evidence

Stated because S4 requires recording what a surviving conclusion beat. Baby's First Years (Gennetian et al., Nature Human Behaviour 8:1514–1529) — an RCT of unconditional cash to mothers of infants — found no statistically significant difference in mothers' paid work or children's time in childcare. Cash at plausible magnitudes has a measured null on precisely the objective the filing's labor_supply_first weighting scores. The builder found this itself and reported it against its own case, which is what an honest steelman looks like.

Overall verdict on Steelman A: partially survives (S4 outcome 2). The filing's recommendation to fund supply is not overturned by this point — a3 and a8 hold top-four places under every weighting in both specifications. What falls is the claim that four designs survive every ranking; the claim that cash and employer instruments rank last on the evidence (they rank last partly on definitional grounds); and any claim that a3 leads, which it does under no weighting in either specification — a8 leads all four on the committed board.

S5 — Scorecard consequences

Two independent sensitivities now exist, both published in scorecard/rationale.md and above:

  1. Under Phase 1's corrected Byrd finding (procedural re-score, computed on the pre-2026-08-07 board): top-4 set unchanged; a8 ties a3 under child_dev_first and leads under labor_supply_first. Superseded as a statement about a3's primacy — the blind re-score has since put a8 first under all four weightings outright, so there is no a3 lead left for this sensitivity to weaken.
  2. Dropping instrument-labelled dimensions: top-4 set changes in three of four weightings — a6 out, a11 in.

Does the recommendation move? Partly. "Fund supply, fund compensation, house standards inside federal spending programs" survives both. "These four architectures, stable under every weighting" does not survive the second. The defensible published claim is: a8 (public option) leads under every weighting on the committed board; a3 (Head Start scaled) holds a top-four place under every weighting in both specifications but leads none of them; and the remainder of the top tier is weighting- and specification-dependent, with one member (a6) an artifact of three dimensions that score instrument type rather than outcome.

S3 / S4 — Steelman C: is universal 0–2 non-parental care desirable at all?

Independently-derived tilt: "pro-universal-coverage on the goal, skeptical on the means. Universal coverage of all 51.2M children 0–12 enters as the definition of the problem." The builder reached the same observation the adjudicator did about the filing's own steelman — it argues the build is more achievable, so "the unfashionable side it identified was optimism about the build, never pessimism about the goal."

C-1. The apparatus cannot express the question — STEELMAN WINS, verified from the repo

Three findings, all checkable without leaving the repository, all verified by the adjudicator:

  1. No dimension measures child development. All fourteen dimensions in scales.md score money, coverage, procedure, federalism, durability, pass-through, participation, preference fit, integrity, distribution, evaluability. None scores an effect on children.
  2. The child_dev_first objective is implemented as coverage. rank.py, verbatim: "child_dev_first": {…, "relief_workforce": 3, "coverage_34": 2, "preference_fit": 2, "evaluability": 2, "coverage_02": 2}. The "child development" weighting scores workforce relief highest (×3) and otherwise rewards covering more children. A design scores better for child development by enrolling more under-2s — which is precisely the contested proposition, entered as an axiom of the objective function rather than tested.
  3. No hypothesis could return "don't do this." adjudication-criteria.md has no supported-branch meaning the goal is wrong.

C-2. The filing's own protocol review demanded this work, named it, and it was not done — STEELMAN WINS

This is the most damaging item in Phase 2, and it is entirely internal. protocol-review.md:

Three literatures named. The filing delivered one. ws07-quality.md presents Tennessee VPK adversarially against Boston — exactly as instructed. Quebec child outcomes and Head Start fade-out were not delivered at all:

The protocol review predicted the omission, named the exact three literatures, the search retrieved the papers immediately, and two of the three were dropped. That is not a sourcing gap; it is a specified deliverable that was not performed and not disclosed — nothing in the deviations log records it.

C-3. The substantive case — PARTIALLY SURVIVES, and the builder said so first

Marshalled at the required tier: BGM (JPE 2008; AEJ:Policy 2019) on persistent non-cognitive harms and crime effects; Havnes–Mogstad (JPubE 2015) showing the same Norwegian reform is positive to the 69th percentile and significantly negative above the 81st; Cornelissen et al. (JPE 2018) finding reverse selection on gains; Fort–Ichino–Zanella (JPE 2020) finding IQ costs at 0–2 rising with family income in a 1:4/1:6-ratio system, which forecloses the "just low quality" defence; Drange–Havnes (JOLE 2019), a real randomized lottery, finding no effect among high-income families.

The builder then reported the evidence that defeats its own strong claim, which is what S3 asks for: Montpetit, Carrer & Beauregard (April 2026) find BGM's behavioural harms "do not translate into lower educational attainment or reduced earnings" and estimate MVPF 2.83; Kottelenberg–Lehrer (JOLE 2017) find gains for disadvantaged single-parent children, inverting the gradient's location; Herbst (JOLE 2017) finds the US Lanham Act — universal — produced +0.36 SD concentrated among the most disadvantaged.

Verdict: partially survives (S4 outcome 2). It does not establish that universal 0–2 care is harmful, and it does not claim to. It establishes that the question is live, that the field is genuinely divided, and that the filing's instrument cannot represent either side.

C-4. The result that matters most — a right answer resting on a wrong reason

The steelman vindicates the filing's 0–2 architecture while dissolving its justification. ws14/ws16 reach "supply-first plus caregiver-choice blend, not centre-led" for 0–2 via revealed preference and supply economics. The child-outcome literature arrives at the same destination by an entirely different route: returns are highest where the home counterfactual is weakest, and kin/home settings preserve near-1:1 ratios. The filing is right about 0–2 for reasons it never examined — which makes the conclusion fragile to precisely the objection it never heard.

C-5. S5 — and an honest refusal

Verified by the adjudicator: setting coverage_02 to weight 0 in child_dev_first — merely declining to assume that covering under-2s benefits children — leaves the top-4 set unchanged (a3 3.89 · a8 3.83 · a6 3.83 · a9 3.56). The recommendation survives agnosticism.

The builder computed that inverting the dimension changes the set (a9 out, a4 in) and then declined to offer it as the correct re-score, because inversion also demotes a11 caregiver-choice — the architecture its own evidence most supports. That refusal is the right call and it sharpens the finding: the scorecard has no dimension capable of carrying this question, and forcing it into coverage_02 corrupts a dimension that means something else. The fix is a new child_effect_0_2 dimension, not a re-weighting. Recorded as a recommendation, not applied — adding dimensions to a published scorecard is the filing's call, not the checker's.


Convergence across independent builds

Steelmen A and C were built in separate contexts, on different targets, with no access to each other. Both independently arrived at the same structural defect from opposite ends: the scorecard's dimensions encode the filing's conclusions. A found three dimensions whose scale anchors name the instrument category; C found that the objective meant to represent child development contains no child-outcome term. Both also independently surfaced the same omission — the Head Start Impact Study — which the filing's own protocol review had named.

Independent convergence on a defect neither was pointed at is the strongest evidence this phase produced.

S3 / S4 — Steelman B: is the constraint regulatory rather than a factor price?

Independently-derived tilt: the filing holds "a factor-price diagnosis of scarcity" — care is scarce because providers are underpaid, so the fix is money routed to compensation. The builder then showed the alternative was excluded by construction, not argued against:

B-1. The headline number is a function of a parameter the filing never examined — STEELMAN WINS

fte_model.py treats staff-child ratios as an exogenous technological constant. Verified by the adjudicator, re-running the committed model with only the ratio parameters changed and every other value at the filing's own settings:

Ratio regime FTEs required Net new Hires/yr
Filing central (published) — r = (4.5, 10.0, 14.0) 2,807k 1,757k 798k
Florida statute (Fla. Stat. § 402.305(4)) — r ≈ (7, 17, 25) 1,727k 677k 501k
Looser still (Mississippi-shaped) 1,623k 573k 473k

The decisive comparison. Halving turnover — the filing's own "arithmetically dominant lever," the thing Part 2 is built around — takes hiring from 798k/yr to 509k/yr. Adopting the ratios Florida's own statute already permits, with no pay rise at all, takes it to 501k/yr. The two levers are equivalent in magnitude; one costs $45.9B/yr and the other costs nothing and needs no federal vehicle.

A second point sharpens it: the model's r02 = 4.5 is tighter than the median state's law and tighter than NAEYC's own recommendation, while the docstring describes the values as "licensing-typical." And the filing praises Florida in Part 5 as a durable multi-decade success while modelling ratios roughly 1.7–1.8× tighter than Florida requires.

B-2. "Binding" fails on the filing's own definition — STEELMAN WINS

research-inquiry.md defines the term: a constraint is binding if relaxing it alone would materially raise achievable output and relaxing the others would not. Relaxing the ratio alone does approximately what relaxing compensation does. The second clause fails on the filing's own test. The word doing the most work in the whitepaper — stamped Binding constraint in Part 2 — is not established.

B-3. What the regulation buys — STEELMAN PARTLY WINS

Marshalled at tier: Hotz & Xiao (AER 101(5) 2011) — one fewer infant per staffer reduces centres in the average market 9.2–10.8%, concentrated in poorer markets; Ali, Herbst & Makridis (IZA DP 14684), the best-identified design either side has — new group-size regulation cuts childcare job postings 5.5–8.4% and lowers maternal LFP 2.0pp, which inverts the filing's model by making tighter ratios reduce childcare employment; Perlman et al. (PLOS ONE 2017) — within permissible ranges, ratio variation shows pooled r = 0.03 with child outcomes, threshold story tested and rejected; Vanderbilt's Prenatal-to-3 Center rates ratios "Needs Further Study." And the filing's own §7 supplies the point: what separated Boston from Tennessee was "curriculum + coaching + pay near parity" — all process quality, not one structural ratio.

B-4. Where the steelman breaks, in its own words

Recorded because S4 requires it, and because the builder volunteered all three:

Verdict on Steelman B: partially survives (S4 outcome 2) — on the load-bearing word

The filing is right that at its assumed ratios the required flow is unprecedented; right about the five-country pattern; right that low pay drives turnover. What does not survive is "binding" as a claim about the world rather than about a parameter. Applied to the record per S4, not footnoted.

S5 consequence

The scorecard is structurally incapable of testing this: it contains no regulatory-reform architecture and no cost-per-slot or regulatory-burden dimension. No re-score can be offered, because there is nothing to re-score — which is itself the finding. Recorded as a recommendation: a future pass needs a regulatory_reform architecture and a cost_per_slot dimension before this question can be scored at all.


Convergence across all three builds

Three cases, three targets, no access to each other. All three land on the same structural fact from different directions: the filing's instrument encodes its conclusions. A found three dimensions anchored on instrument type; B found no dimension capable of registering regulatory cost and no architecture that alters it; C found the child-development objective built entirely from coverage terms. And all three found commissioned work that was never delivered and never logged — Quebec child outcomes and Head Start fade-out (C), facilities codes and zoning (B), QRIS evaluation evidence (both).

The pattern, stated once: GBMT-1's protocol was better than its execution, and the gap between them is not recorded anywhere in the filing. That is the finding this phase exists to produce, and it is more useful than any single corrected number.

← All Childcare research documents Sources digest Read the whitepaper