GUBMENTPlain talk · policy frontier
Filings / Childcare / Sources / Phase 2 Steelman — GBMT-1 Childcare —
GBMT-1 · Research record · No. 1

Phase 2 Steelman — GBMT-1 Childcare — Target: cash to families / demand-side / employer channel

childcare/research/steelman-cash.md
This is a working research document from the childcare filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 1, not a summary of it.

Protocol: method/verification-protocol.md, Phase 2 (S1–S5). Built 2026-08-04. Independence note (S1/V1): built without reading the Phase 1 verification log's verdicts, without reading red-team.md, and without any external party's account of this filing's weaknesses. Every source below was fetched in this session; no figure appears here that I did not read inside the fetched document itself. Where I have only a search summary, it is marked UNVERIFIED.


1. S2 tilt statement — written before any steelman evidence was gathered

Derived only from the filing's own text (site/childcare/index.html, site/childcare/sources/index.html, childcare/research/adjudication-criteria.md, ws14-architectures-v0.md, ws08-demand.md, ws15-precedents.md, scorecard/{scales,scores,rationale,rank}).

Which direction it leans

The filing leans, unambiguously, toward publicly financed supply of formal, center-anchored, credentialed childcare, with the money entering on the provider side. Its abstract states the conclusion as an ordering — "pay the workforce first, build supply second, extend benefits third" — which places every demand-side instrument third in time and contingent in principle (§7: "extend demand-side eligibility only as staffed capacity arrives"). Demand-side money is framed throughout not as an instrument with its own benefits but as a risk to be sequenced around.

Which conclusion the evidence was marshalled toward

Toward "the binding constraint is the childcare workforce." ws15 calls this "the protocol's central working hypothesis" and reports that it "now has support from five independent national designs." Five out of five heterogeneous national cases confirming a single hypothesis is a pattern consistent with a frame that preceded the survey: each case is read for how it confirms (Quebec = queues, Germany = litigation, Australia = price, Canada = shortage, DC = wages work), never for whether it disconfirms.

Which architecture won

scorecard/rankings.txt: a3 Head-Start-scaled, a8 public option, a6 supply-first, a9 fallback ladder — top four under all four weightings, no flips. All four are supply-side public-provision or public-purchase designs. a5 demand-side allowance ranks 11th of 12 under all four weightings; a10 federalized Tri-Share (the employer channel) ranks 12th under three of four; the unconditioned cash_comparator ranks 8th–10th. The public page compresses this to: "The two most fashionable designs — unconditioned cash and employer cost-splitting — rank last."

Whose framing it adopted

The early-childhood-education professional field's. Specifically:

The position it argued against, scored down, or never seriously entertained

  1. Cash / demand-side to families (a5, a11, cash_comparator) — argued against on exactly two cautionary cases (Australia fee inflation; Netherlands clawback) and one theoretical channel (pass-through into inelastic supply).
  2. The employer channel (a10) — dismissed on a single data point (Michigan Tri-Share's ~$14M cumulative savings) plus an unresolved hypothesis (H12.3) that adjudication-criteria.md had pre-registered as expected to come back indeterminate.
  3. Never seriously entertained: that the supply-side designs ranked 1–4 might not produce the child-development benefits that justify their cost.

steelman-feasibility.md steelmans feasibility of the recommendation, not a rival instrument. No document in the record steelmans an alternative architecture. M7 is unmet as to instrument choice.


2. The steelman: the ranking is wrong

The filing's ranking rests on a single implicit premise — that a formal childcare slot is the unit of value, so an instrument's merit is how reliably it produces slots. Once that premise is stated, it can be tested. Five independent lines of evidence say it is wrong, or at least that it is a premise the filing adopted without ever testing it.


Point 1 — Three of the fourteen scorecard dimensions define the answer, and removing exactly those three breaks the filing's headline finding

This is repo-internal and needs no external source. scorecard/scales.md anchors:

Dimension "1 means" "5 means"
passthrough_risk "Pure demand-side into inelastic supply" "Supply-side operating grants / direct provision"
integrity_risk "Netherlands-shaped: mass demand-side clawback exposure" "Direct provision / small-N grantee audit surface"
relief_supply "Builds no capacity; pure transfer" "Builds new capacity incl. rural/nontraditional"

The endpoints of the first two scales are the names of the two instrument classes being compared. A demand-side design cannot score above 1 on passthrough_risk and a direct-provision design cannot score below 5, regardless of any evidence. These are not dimensions; they are the architecture label re-entered as data, three times, in a 14-cell average. relief_supply is the same defect in milder form: it scores an instrument on whether it is the kind of instrument that builds capacity.

To be fair to the filing: this is not uniform. procedural, federalism, participation_risk and preference_fit are genuinely open, and cash scores 4–5 on all four. The claim is narrower and sharper: exactly three dimensions are definitionally closed, and closing them is what produces the headline result.

Sensitivity (S5). I reproduced rank.py exactly (baseline output matches rankings.txt line for line), then re-ran it dropping only those three dimensions:

Filing (14 dims) Sensitivity A (11 dims)
Top-4 under all weightings a3, a6, a8, a9 — "no flips" a3, a8, a9 only — a6 drops out
a6_supply_first 3rd, 3rd, 3rd, 4th 5th, 5th, 4th, 5th
a11_caregiver_choice 6th, 7th, 5th, 5th 4th, 4th, 5th, 3rd
cash_comparator 8th, 10th, 10th, 8th 8th, 10th, 11th, 6th
a5_demand_allowance 11th ×4 9th, 9th, 9th, 8th

Supply-first — which ws14-architectures-v0.md calls "the session's biggest winner" — leaves the stable top-4 the moment the three definitional dimensions come out, and the architecture that takes its place under three of four weightings is the caregiver-choice allowance, i.e. cash for the 0–2 band. The rationale file's own claim that "rank stability of this degree means the recommendation does not depend on resolving the §1.6 objective question — the strongest kind of result the method can produce" does not survive the removal of dimensions whose anchor text names the winner.

Two absences compound this. The scorecard has no cost-per-child-served dimension and no child-outcome-evidence dimension. Both omissions run the same direction: the supply-side designs are the expensive ones (Point 3) and the ones with null RCTs (Point 2).

Reproducible script: scratchpad/sensitivity.py (run as python3 sensitivity.py scores.csv, with scores.csv copied unmodified from childcare/research/scorecard/). Its baseline block reproduces rankings.txt line for line; Sensitivity A's stable top-4 prints as ['a3_headstart_scaled', 'a8_public_option', 'a9_fallback_ladder'].


Point 2 — The filing's top-ranked architectures are the ones with null or negative randomized evidence; cash is the one with positive quasi-experimental evidence. The filing cites neither literature.

This is the substantive core. The filing ranks a3 (Head Start scaled) first and never mentions that Head Start has a national randomized controlled trial commissioned by Congress.

a3's own RCT — HHS/ACF/OPRE Report 2012-45, Third Grade Follow-up to the Head Start Impact Study: Final Report (nationally representative sample of Head Start programs; children randomly assigned to access). Verbatim from the report PDF, fetched and extracted this session:

"Looking across the full study period, from the beginning of Head Start through 3rd grade, the evidence is clear that access to Head Start improved children's preschool outcomes across developmental domains, but had few impacts on children in kindergarten through 3rd grade."

"In summary, there were initial positive impacts from having access to Head Start, but by the end of 3rd grade there were very few impacts found for either cohort in any of the four domains of cognitive, social-emotional, health and parenting practices. The few impacts that were found did not show a clear pattern of favorable or unfavorable impacts for children."

The report also records that at 3rd grade "there was suggestive evidence of an unfavorable impact — the parents of the Head Start group children reported a significantly lower child grade promotion rate than the parents of the non-Head Start group children." Source: https://acf.gov/sites/default/files/documents/opre/head_start_report_0.pdf (Puma, Bell, Cook, Heid, Broene, Jenkins, Mashburn, Downer, OPRE Report 2012-45, October 2012).

Corroborating, at scale, on the successor design. Durkin, Lipsey, Farran & Wiesen, "Effects of a statewide pre-kindergarten program on children's achievement and behavior through sixth grade," Developmental Psychology 58(3): 470–484 (2022) — an RCT, n = 2,990 low-income children randomised to admission offers vs. waitlist at oversubscribed Tennessee VPK sites. From the PubMed record of the article (PMID 35007113) and its published correction (PMID 35759004, Dev Psychol 58(7):1385):

"children randomly assigned to attend pre-K had lower state achievement test scores in third through sixth grades than control children, with the strongest negative effects in sixth grade. A negative effect was also found for disciplinary infractions, attendance, and receipt of special education services."

(The correction is honest disclosure, not a retraction: it fixed reverse-coded treatment indicators in two supplemental logistic tables; "the magnitude and p values were correct.")

The filing's own comparative case, read at its own literature. The filing uses Quebec five times (waitlists, staffing). Quebec is also the most-studied universal-childcare natural experiment in the world, and what that literature found is not in the record:

So Quebec — the filing's own exemplar — delivers a crowd-out finding (one-third of new formal use displaced informal care families were already using) and a negative child-outcome finding, at the exact age band the filing's 0–2 supply build-out targets. The filing reports only the waitlist.

Now the other side of the ledger. Cash has the positive evidence, in the same country, from the same journals.

The comparison that should have been in the filing. In Canada, across overlapping periods, econometricians using the same methods and publishing in the same journals found that the supply-side universal childcare instrument produced worse non-cognitive outcomes persisting into adulthood, and the unconditioned child cash benefit instrument produced better test scores and better maternal and child mental health. The filing scored twelve architectures under a "child-development-first" weighting without touching either result.


Point 3 — The cost per unit is not close, and the ASPE data says the cheap version is the one that looks like cash

HHS/ASPE, Head Start Spending Per Slot Varies Widely Across Grants, Issue Brief, February 2, 2026 (fetched and extracted this session):

"The median amount grantees spend per slot is 20, 294inEarlyHeadStart * *and * *14,532 in Head Start Preschool." "In Early Head Start, the largest and most persistent differences in spending per slot are associated with service delivery setting. Grants that provide only home-based services spend about 33 percent less per slot compared to grants that provide only center-based services."

Source: https://aspe.hhs.gov/sites/default/files/documents/6b4fa8b4c6e481fdb83cae736c632425/Head%20Start%20Spending%20Per%20Slot%20Brief_Final.pdf

Set that against the cash instrument that has actually been randomised in the United States: Baby's First Years assigned mothers 333/month— * *3,996/year** (Sperber et al., JAMA Network Open 6(9): e2335237, 2023; PubMed record fetched). One median Early Head Start slot costs the same as 5.1 child-years of the Baby's First Years cash treatment.

Two consequences the scorecard has no cell for:

  1. Any budget-constrained comparison of a3/a6/a8 against cash must be run per dollar, not per design. The scorecard's 1–5 scales are unit-free and therefore silently assume equal cost.
  2. ASPE's own decomposition says the home-based delivery mode — the one that resembles what cash buys, and the one NSECE says families with under-3s actually use — is a third cheaper per slot than the center-based mode the filing's architecture is built around.

Corroborating the appropriation scale: FY2023 Head Start (incl. Early Head Start–Child Care Partnership) was funded at $11,589,715,163 for 778,420 funded slots, per ACF's own Head Start Program Facts: Fiscal Year 2023 (https://headstart.gov/sites/default/files/pdf/hs-program-fact-sheet-2023.pdf).


Point 4 — The filing's own §8 revealed-preference finding is stronger than the filing states, and it points away from a center build-out

I re-fetched the primary source rather than relying on ws08's summary. ACF/OPRE, Children's Participation in Child Care and Early Education in 2012 and 2019: Counts and Characteristics (NSECE Chartbook, May 2023), Appendix Table 1b and Exhibit 3 (https://acf.gov/sites/default/files/documents/opre/NSECE%20HH%20Chartbook%20CCEE%20Participation__toOPRE_05302023_508_compliant.pdf):

The filing reports the ws08 finding as "under-3s' modal arrangement is unpaid individual care" and then scores a11 coverage_02 = 5, preference_fit = 5 and calls it "the only architecture matching revealed 0–2 preference" — before ranking it 5th–7th. But the sharper number is the base rate: a center-based national offer for the 0–2 band is a build-out aimed at the 13.4% of under-3s who currently use centers, financed at $20,294 per slot, for a population where a majority uses no regular non-parental care. Cash is the only instrument in the twelve that is neutral between the arrangement 26.5% actually use, the arrangement 13.4% use, and the choice 51.5% have made.

The filing's honest counter — that non-use may be rationing rather than preference — is real and I grant it in §3. But it is an assumption the filing needs and never tests, and it is doing enormous work: it converts a revealed-preference finding into a latent-demand finding without evidence.


Point 5 — Australia is the wrong cautionary tale, and the primary statistics say the opposite of what the filing says

The filing's entire empirical case against demand-side instruments outside theory is Australia. Its source (ws15-precedents.md) is a Guardian article dated 14 October 2020 via pressreader — press tier, one year, three years before the reform that matters.

Australia then ran the largest demand-side subsidy increase in its history (Cheaper Child Care, effective 10 July 2023). The Australian Bureau of Statistics — the national statistical agency, i.e. the top available tier for this claim — reported the result (media release, CPI rose 1.2 per cent in the September 2023 quarter, fetched from abs.gov.au):

"Child care fell 13.2 per cent, and was the largest contributing fall this quarter. Changes to the Child Care Subsidy raised the amount of subsidy received for over a million families and came into effect on 10 July 2023. 'This change reduced out of pocket costs for households, more than offsetting child care fee increases this quarter. Without the changes to the Subsidy, child care would have increased 6.7 per cent,' Ms Marquardt said."

And then it got better for the demand-side case, not worse. ABS Media Statement, Forthcoming correction to child care costs in the Consumer Price Index, released 19/11/2024 (https://www.abs.gov.au/media-centre/media-statements/forthcoming-correction-child-care-costs-consumer-price-index):

"The ABS has recently identified it made errors in estimating the impact of the Government's reforms to increase the rate of the Child Care Subsidy (CCS) … As a result, consumers' out-of-pocket child care costs were overstated in the Consumer Price Index (CPI) from the September quarter 2023 onwards. In the most recent CPI publication for the September quarter 2024, the published Child care index was 5.8 per cent (or 9.5 index points) higher than it should have been."

So: the demand-side instrument delivered a 13.2% one-quarter reduction in what families actually paid, against a counterfactual 6.7% fee increase — and the official statistics understated the pass-through to families by a further 5.8%. That is roughly 60% of a large subsidy increase reaching households inside one quarter, in the very market the filing describes as the archetype of demand-side capture.

This does not refute pass-through as a mechanism. It refutes the filing's use of Australia as evidence that pass-through defeats demand-side instruments, and it makes a5 passthrough_risk = 1 — a cell whose stated basis in rationale.md is "Australia fee inflation (ws15)" — unsupported at the tier the Verification Protocol requires. A national statistical agency's price index beats a 2020 newspaper clipping, and it says the opposite.


Point 6 — The employer channel is the only one of the twelve that has already cleared the filing's own procedural test

The filing's Part 4 is an argument that Senate procedure is a first-order design constraint and that what survives reconciliation is money, not mandates. Apply that test to a10.

26 U.S.C. §45F, current text via uscode.house.gov (statutory tier — the filing's record cites this at the BPC/IRS tier in ws12-private-capital.md):

"(a) … (1) 40 percent (50 percent in the case of an eligible small business) of the qualified child care expenditures … (b)(1) The credit allowable … shall not exceed 500, 000(600,000 in the case of an eligible small business). (2) … [indexed for taxable years beginning after 2026]"

Amendment note: "(Added Pub. L. 107–16 … amended … Pub. L. 119–21, title VII, §70401(a)–(f), July 4, 2025, 139 Stat. 212, 213.)" — §70401(a) substituted 40%/50% for 25%; §70401(b) replaced the $150,000 cap.

Three things follow that the scorecard does not reflect:

  1. The employer channel is the only childcare architecture in the twelve that was actually enacted federally, and it was enacted in a reconciliation act. The filing scores a10 procedural = 3 ("reconcilable with major strips"). The statute says clean fit, already done, 2.7× the credit rate and 3.3× the cap.
  2. §70401(d) added to §45F(c)(1)(A)(iii) the words "or under a contract with an intermediate entity that contracts with one or more qualified child care facilities to provide such child care services." That intermediate-entity path is the exact legal mechanism a Tri-Share pool needs, and it did not exist before July 2025. The record notes the rate and cap changes but not this one.
  3. §45F(c)(1)(A)(ii) already treats as a qualified expenditure "the operating costs of a qualified child care facility of the taxpayer, including costs related to the training of employees, to scholarship programs, and to the providing of increased compensation to employees with higher levels of child care training." The employer channel therefore already funds compensation — the filing's own first-order goal — and does so through the instrument the filing's procedural analysis says is the survivable one.

The filing ranks this last on evidence that adjudication-criteria.md pre-registered as expected to be indeterminate (H12.3: "Indeterminate otherwise — expected; this may not resolve from US data") plus one state pilot's savings figure. A 12th-place ranking is not a permissible output of an indeterminate adjudication.


Point 7 — What cash does not require: fifty state governments, 800,000 hires a year, and a credential pipeline that issues 40,000

The filing's own arithmetic is the strongest argument for the instrument it ranks last. It reports a requirement of ~2.8M FTEs against ~1.10M employed, ~800,000 hires/year of which 579,000 are replacement, against a credential pipeline issuing ~40,000/year — and states in its own honesty box that "nobody has tried fair wages at scale" and that the workforce ramp has no precedent.

Cash to families has a delivery mechanism that already exists, ran at national scale within six months of enactment in 2021, and produced the measured result. Census Bureau, Poverty in the United States: 2022, Report P60-280 (Shrider & Creamer), fetched from census.gov:

"The SPM child poverty rate more than doubled, from 5.2 percent in 2021 to 12.4 percent in 2022." "Refundable tax credits moved 6.4 million people out of SPM poverty [in 2022], down from 9.6 million people in 2021."

Whatever one thinks of the CTC expansion's design, the on-off pattern is the cleanest evidence in American social policy that a demand-side transfer can be turned on nationally, reach households at scale, and be measured — in a domain where the filing's preferred instrument requires a workforce ramp it concedes is unprecedented and a state-participation assumption CBO already scores as failing.


The steelman in one paragraph

The filing chose slots as the unit of value, then built a scorecard three of whose fourteen dimensions score architectures on whether they produce slots, then found that slot-producing architectures win. When those three dimensions come out, its "biggest winner" leaves the stable top four and cash for the 0–2 band enters it. The architectures it ranks first are the ones whose randomized evaluations found few impacts by third grade (Head Start Impact Study) or negative impacts through sixth (Tennessee VPK), and whose closest international analogue produced persistent non-cognitive deficits, worse adult health and higher crime (Quebec, BGM 2019) — while displacing informal care one-third of families were already using (BGM 2008). The instrument it ranks last is the one with positive quasi-experimental child-outcome findings in three separate top-field journals (Milligan–Stabile 2011; Dahl–Lochner 2012; Hoynes–Miller–Simon 2015), costs roughly a fifth as much per child-year as the design it ranks first ($3,996 vs. $20,294), is neutral across the arrangements 51.5% of under-3 families have chosen, and is the only one already delivered at national scale and measured by the Census Bureau. Its single cautionary case, Australia, is a 2020 newspaper clipping that the Australian Bureau of Statistics contradicted in 2023 and then corrected further in the demand side's favour in 2024. And the employer channel it ranks twelfth is the only one of the twelve that has actually passed Congress, in a reconciliation act, under the filing's own procedural theory.


3. Honest self-assessment

Where this case is strong

Where it breaks

The filing's best rebuttal

Stated as strongly as I can make it, because it is good:

"The steelman conflates three different questions. (1) On child development: preschool's best-identified modern American evidence is not HSIS (a 2002-cohort study of programs since reformed) but the Boston lottery — Gray-Lobe, Pathak & Walters, 'The Long-Term Effects of Universal Preschool in Boston' (NBER w28756) — which found preschool 'boosts college attendance, as well as SAT test-taking and high school graduation' and 'decreases several disciplinary measures including juvenile incarceration.' Modern, well-run, publicly delivered preschool works; Tennessee's implementation is a quality finding, not an architecture finding, and the filing's whole thesis is that quality is a compensation problem. (2) On labour supply: the steelman's own headline RCT found cash moved neither maternal employment nor childcare use. Our objective function includes maternal employment; cash has a null there and supply-side has a measured elasticity. (3) On Australia: a one-quarter 13.2% out-of-pocket fall is the mechanical arithmetic of a subsidy step-change, and the ABS release itself records that underlying fees rose 6.7% in the same quarter — which is the pass-through claim, not a refutation of it. The scorecard's passthrough_risk anchor is not circular; it is an empirical judgement, held since Quebec and Australia, that where supply is inelastic, demand-side money becomes provider revenue. Finally, cash and supply are not rivals in our design: we recommend caregiver-choice for 0–2 alongside supply-first, and we sequence demand-side expansion after capacity precisely so it buys care rather than price."

The strongest part of that rebuttal is (2) — the BFY null — and the Boston lottery result in (1), both of which I have verified. The weakest part is the defence of passthrough_risk: an empirical judgement belongs in a cell, defended by evidence, not in a scale anchor where it applies automatically to every architecture in its class and cannot be scored against.

Adjudication under S4

The steelman partially survives. It does not defeat the filing's core conclusion — that the workforce is the binding constraint on formal capacity, and that supply-side money suppresses prices where demand-side money may not. The BFY null on maternal employment and childcare use is a real defeat for the strong form of the cash case.

But it defeats three specific claims:

  1. "Four designs occupy the top ranks under all weightings" — withdraw or qualify. The stability is an artefact of three definitional dimensions; supply-first leaves the stable set when they are removed, and caregiver-choice enters it under three of four weightings. Publish the sensitivity.
  2. a5 passthrough_risk = 1 and the sentence "demand-side subsidies (Australia) still inflated fees" — CORRECT. The ABS, at statistical-agency tier, records a 13.2% out-of-pocket fall from the 2023 CCS increase and a subsequent correction finding the fall was understated by a further 5.8%. The filing's basis is a 2020 press clipping.
  3. "The two most fashionable designs … rank last" — withdraw as to the employer channel. a10's ranking rests on a hypothesis the filing pre-registered as expected-indeterminate, and its procedural = 3 is contradicted by Pub. L. 119–21 §70401, which enacted the channel's expansion through reconciliation on July 4, 2025.

And it identifies one omission large enough to be its own finding: a filing that scores twelve architectures under a "child-development-first" weighting, and ranks a scaled Head Start model first, without citing the Head Start Impact Study, the Tennessee VPK RCT, the Boston lottery study, or the Quebec child-outcome literature, has not evaluated child development. That weighting should either be scored against that evidence or removed.

← All Childcare research documents Sources digest Read the whitepaper