Protocol & scope
United States childcare, birth through age 12 — federal plus 50 states, DC, and territories. Cost and financing, workforce, physical supply, quality regulation, delivery architecture and federalism, demand, market response, political economy, legislative and legal mechanics, private capital and the employer channel, civil society, candidate implementation architectures, precedents, and sequencing. Imports the same discipline as every gubment filing: two-source rule, root-tracing for every claim, pre-registered hypotheses with written adjudication criteria fixed before evidence collection, and a deviations log that stays visible rather than getting edited away.
Baseline plus 14 workstreams executed across three research passes on 2026-08-03. Phase 0 alone moved two priors before full execution even began: Canada's national rollout was already underdelivering on the workforce constraint this protocol treats as primary (194,000 of 284,000 targeted new spaces delivered), and CBO's own Build Back Better score turned out to already assume substantial state non-participation as its base case, not a tail risk. A formal architecture scorecard followed (11 architectures plus a cash comparator, scored on 14 anchored dimensions under four objective weightings), then a red-team pass applying three confidence downgrades, and a symmetric feasibility steelman built to the same evidentiary standard as the protocol's uniformly skeptical hypotheses.
Anchor table — priors, stated before evidence, verified after
Every anchor below was written down as an unverified guess before research began, so it could be broken. ★ rows were verified in Phase 0; the rest were closed (or left queued) across passes 2–3. Eight rows remain blank by rule — not guessed, not smoothed over.
| # | Anchor (unverified prior) | Verified value & delta |
|---|---|---|
| 1 ★ | BBB childcare/pre-K title ≈ $400B / 6yr as House-passed (Nov 2021) | CBO scored the childcare+pre-K provisions at +$381.5B (2022–2031), not the $400B bill line item — and the score itself assumes substantial state non-participation, treating the federalism problem as the base case rather than a tail risk. |
| 2 ★ | ARPA stabilization ≈ $39B, expired 2023-09-30 | Confirmed as two tranches, not one cliff: $24B stabilization funding expired 2023-09-30; a further $15B CCDF supplemental wound down through September 2024. |
| 3 ★ | Childcare cost is ⅔–⅘ labor | Treasury puts wages at 50–60% of expenses on US averages (higher for infant care); center personnel costs are commonly cited at 70–80%. Widen the range to 50–80% rather than the protocol's narrower floor. |
| 4 | ≈75 employment-relevant non-school weekdays per year | Not verified in this pass — queued for future §2 baseline execution. |
| 5 ★ | Additional workforce needed: several hundred thousand to ~1M | The FTE model puts steady-state need at 2.8M FTEs (range 1.9–3.9M) against a 2020 pre-pandemic baseline of ~1.05M industry jobs (current employment is closer to 1.10M) — net new ≈1.76M against the model's baseline. The prior understated the requirement by roughly 2×. |
| 6 | Tax-increase supermajority requirements in ≈12 states | 17 states, not ~12, require legislative supermajorities to raise taxes — the state-revenue path is harder than the prior assumed. |
| 7 | Ballot initiative available in ≈half the states | Not yet verified — queued. |
| 8 | CCDBG authorization lapsed after the 2014 reauthorization period | Confirmed: the 2014 reauthorization (P.L. 113-186) covered FY2015–FY2020 and expired; the program has run on appropriations alone since. |
| 9 | FRA 2023 discretionary caps covered FY24–25 | Not yet verified — queued. |
| 10 | Dependent care FSA cap nominally near-frozen for decades (brief ARPA-era increase) | Confirmed and now superseded: the $5,000 cap dates to 1986 and was never indexed — but it was raised to $7,500 effective 2026 (OBBBA, signed July 2025). The employer/tax channel is now expanding on two tracks at once (with §45F), making the crowd-out hypothesis a live, trackable experiment rather than a hypothetical. |
| 11 | §45F persistently under-claimed relative to authorization | Not yet verified as an anchor row — queued. |
| 12 ★ | Georgia universal pre-K mid-1990s; Oklahoma late 1990s; Florida mid-2000s | Confirmed. Georgia enacted the nation's first universal pre-K in 1995 (lottery-funded, ~55% of 4-year-olds, 60% of classrooms private); Oklahoma followed in 1998 via the school-funding formula, reaching over 70% of 4-year-olds — the nation's highest; Florida's ballot-approved program has operated since 2005. Exact Florida amendment-vs-launch dates still to be pinned down. |
| 13 ★ | Canada CWELCC: 2021 agreements, $10/day target | Targets are being missed: only 194,000 of 284,000 targeted new spaces delivered nationally (Sept 2025); Ontario fees average ≈$19/day, not $10, and the $10 deadline has been pushed to December 2026; Ontario alone is short up to 10,000 early childhood educators; 57% of new spaces are for-profit, with an active fight over lifting the for-profit cap. Live confirmation that workforce, not money, binds the ramp. |
| 14 | Netherlands scandal: thousands of families falsely accused; government fell (2021) | Confirmed and worse than stated: an estimated 26,000–35,000 families were wrongly ordered to repay benefits, ethnic profiling was documented, and more than 2,000 children were removed into custody; the government resigned January 2021. |
| 15 | NYC pre-K: tens of thousands of seats in ≈2 years | Confirmed: the funded expansion moved from 20,000 to 53,000 full-day seats in 2014, with more than 60,000 children enrolled within two years — the clearest domestic proof that fast ramp is achievable. |
| 16 | Background-check backlogs of months post-2014 CCDBG in some states | Not yet verified — queued. |
| 17 | Michigan Tri-Share: ≈equal three-way split; Kentucky employer match | Not yet verified as an anchor row — queued. |
| 18 | Quebec reform 1997; Germany slot entitlement from age 1 in 2013 | Not yet verified as an anchor row — queued. |
| 19 | 1971 CCDA passed both chambers, vetoed; CCDBG enacted 1990 | Confirmed via primary source: the bill was vetoed December 10, 1971. The veto message (drafted by Pat Buchanan) framed federal childcare as favoring "communal approaches to child rearing over against the family-centered approach" with "family-weakening implications" — language that still structures opposition framing today. |
| 20 | Head Start since 1965; Lanham centers 1943–46 | Not yet verified — queued. |
Workstream findings
Fourteen workstreams, each executed as a first pass with a full source register and two-source-rule discipline. Full writeups (source-by-source, with hypothesis calls) are in each workstream's own findings file in the repository.
§3 · Cost & financingEvery collected cost estimate excludes school-age care, and labor is the dominant cost share
Four collected estimates (CBO's BBB score at $381.5B, Warren-style plans at ~$70B/year — the source gives no 10-year total — BPC's cost-of-inaction range, average family price) are all scoped to ages 0–5 — none include school-age or summer care, so the full 0–12 cost is understated in every figure circulating. Labor share is confirmed at 50–80% of costs, meaning the §4 wage findings dominate any steady-state estimate. School-age is 64% of the population and absent from every collected model.
§4 · WorkforceThe credentialing pipeline can carry roughly a quarter of the hires needed — retention, not credentialing, is the fast lever
Childcare workers earn a median $15.41/hr, among the lowest-paid occupations BLS tracks, with turnover running ~64% above a typical job on average (65% in 2022). The FTE model requires ~798k hires/year at ramp against a CDA credentialing pipeline producing over 50,000/year — about 23% of growth hiring alone. DC's Pay Equity Fund shows a large turnover gap for centers that took the funding (37% vs 51%), though the finding is cross-sectional, not yet a measured before/after decline; wage parity is ~40% of incremental cost measured against current private spending — a figure that moves with the denominator, and stops short of a majority.
§5 · FacilitiesChildcare deserts are diverging by geography, and licensed-capacity counts overstate real supply everywhere
46% of children under 6 live in a licensed-care desert nationally; remote-rural deserts reached 70% in 2025, while rural areas overall sit at 65.6% — essentially flat against ~66% in 2018 — as the broader post-pandemic recovery was non-rural. State rates range from Alaska's 96% to DC's 5%. Because even licensed capacity in chains sits at only ~71% staffed occupancy, desert measures built on licensed slots understate the true shortage everywhere, not just in deserts. CDFI financing is one facilities-finance channel worth exploring outside conventional bank credit, though how much volume it actually deploys is not yet quantified.
§6 · Delivery, federalism & integrityCBO already scores state non-participation as the base case, not a tail risk
CBO's Build Back Better estimate embedded substantial state non-participation as its baseline assumption — the same holdout dynamic Medicaid expansion produced — strengthening the case that a federal fallback isn't insurance but the difference between a national and a blue-state program. The March 2024 CCDF final rule's co-pay cap, prospective-payment, and direct-services provisions were rescinded by HHS effective July 13, 2026 — three weeks before this filing — after a Senate resolution to restore them failed 47–52 on July 30, 2026; the admin-only path has more headroom than the protocol originally assumed, not less. Germany's enforceable entitlement shows courts can force payment, not supply.
§7 · Quality & the cost–quality frontierTennessee's pre-K RCT found negative effects by 3rd grade; Boston's found real attainment gains — both survive scrutiny
Tennessee VPK's randomized trial found achievement gains gone by kindergarten, negative test-score effects by 3rd grade (worse by 6th), and more discipline and special-ed placement. Boston's lottery study found no persistent test-score effects but real gains in high school graduation, college attendance, and college completion. The synthesis: quality parameters plausibly determine the sign of child effects, not just the size — which raises the stakes on how quality provisions survive Senate procedure at all.
§8 · DemandUnder-3s' modal arrangement is unpaid individual care; nearly 4 in 5 middle-schoolers who want afterschool can't get it
NSECE data confirms the protocol's prior: under-3s are most commonly in unpaid individual care (26.5%), while 3–5s are mostly center-based (42.1%) — a center-only buildout misreads the 0–3 market. About 40% of households with under-5s use no regular nonparental care. Cost burden concentrates just above current eligibility lines: near-poor households (100–200% FPL) face the worst burden while most sub-poverty households pay $0. School-age unmet demand hit an all-time high — 5.2 million of 6.7 million middle-schoolers whose parents want an afterschool program can't get one.
§9 · Market response & price effectsSupply-side grants suppressed prices at $24B scale; demand-side subsidies do the opposite
The ARPA stabilization program is the cleanest US natural experiment on how public money moves this market: costs fell while $24B flowed to 220,000 providers, then prices resumed rising and access deteriorated within two quarters of expiry. Pass-through turns out to be instrument-dependent, not uniform — supply-side grants suppressed prices while demand-side subsidies (Australia) still inflated fees. Chain occupancy at ~71% adds a geography wrinkle: served metro markets have short-run slack; desert markets do not.
§10 · Political economyBBB's childcare title died with its vehicle, not its content — and red-state pre-K has proven durable for decades
The 2021 BBB childcare/pre-K title passed the House and died in the Senate with the whole bill's collapse, not from a childcare-specific defeat; the National Head Start Association actually endorsed BBB's structure. Georgia, Oklahoma, and Florida's universal pre-K programs have survived 20–30 years across partisan turnover. DC's Pay Equity Fund, by contrast, remains annually contested despite proven results — durability requires dedicated revenue, not just evidence.
§11 · Legislative mechanicsReconciliation kills mandates, not money — the $15 minimum-wage fight is the closest on-point test
The Senate parliamentarian advised in February 2021 that the $15 minimum wage did not qualify for the American Rescue Plan on "merely incidental" grounds; Senators dropped it before floor consideration, and a March 2021 amendment to restore it fell to a sustained point of order on the same ground. It's the closest on-point test for any childcare wage floor: compensation requirements face the same objection, but compensation funding is budgetary and generally survives — though CRS cautions every determination is case-specific, and money-moving provisions can still fall to the rule's outyear-deficit test. Quality standards fare best housed inside federal spending programs rather than imposed as market mandates.
§12 · Private capital & the employer channelThe two largest childcare chains are shrinking, not expanding, into the shortage
KinderCare's occupancy fell to ~71% even as revenue grew through price increases, and Bright Horizons closed a net 90 centers between Dec 2022 and Mar 2026 (guiding to a net −25 to −30 for 2025 alone) — both are rationalizing, not building toward the shortage. The §45F employer tax credit was confirmed nearly dead as historically structured (a couple hundred claims, under $20M combined) before a 2026 expansion; Michigan's Tri-Share program shows real but pilot-scale results (~$14M cumulative savings statewide over roughly five years). Nothing found contradicts the hypothesis that private capital can't reduce the core cost of care itself.
§13 · Civil society & coalitionsThe coalition already pre-negotiated a workforce framework; BBB's title fell with its vehicle, not its own fault lines
The 15-organization "Unifying Framework for the ECE Profession" (NAEYC-convened, with NEA, NHSA, SEIU and ZERO TO THREE among signatories) is a genuine pre-negotiated consensus on credentials, pathways, and compensation. Fault lines are real (credentials-without-compensation, a school-based-versus-FCC union-structure split), but BBB's title cleared the House with the coalition intact and died with the whole bill's vehicle, not from an internal split.
§14 · Architecture scorecardThree architectures are stable across every weighting; the fourth seat flips
The formal scorecard — 11 candidate architectures plus a cash comparator, scored on 14 anchored dimensions — was independently re-scored blind on 2026-08-07 (47 cells moved). Three designs occupy the top four under all weightings: public option, Head Start scaled to universal, and K–12 extension. Supply-first holds the fourth seat under three weightings; caregiver-choice takes it under equity. Federal fallback keeps the holdout-proofing cell but leaves the stable set. The shared thread among the leaders is supply-side or direct-provision money and standards housed inside federal spending programs. Demand-side allowance and federalized Tri-Share rank last under every weighting; cash is mid-table.
§15 · PrecedentsFive independent national designs converge on the same binding constraint: educators
Quebec's 25-year-old program has a permanent waitlist (30,688 children as of May 2025) that operators attribute to staffing, not funding. Germany's enforceable legal entitlement to a childcare slot produces litigation and damages payments in staff-short cities, not more slots. Australia's demand-side subsidy saw fees outrun subsidies, prompting the government to add supply-side conditions. Every design that ignored the workforce constraint converted new money into queues, litigation, or price — never capacity.
§16 · SequencingVisible-first where capacity exists, supply-first where it doesn't — dedicated revenue throughout, no sunsets
The draft sequence funds supply and compensation first (Byrd-survivable because it's money, not mandates), with demand-side eligibility expansion only after capacity exists — inverting the failure order seen in Quebec and Australia. School-age sequences first among segments (64% of children, existing buildings, a proven red-state funding-formula precedent), then 3–4 consolidation, with 0–2 last via a supply-first-plus-caregiver-choice blend. A red-team amendment reframed the sequence after GA/OK's demand-first success at the pre-K tier cut against a purely supply-first political strategy.
Deviations log
Every departure from the protocol as originally written, logged with its reason and effect on findings — per the method, this stays part of the record, not an appendix to it. 14 entries; freeze point is the protocol as of commit b12c6fc (2026-08-03).
| # | Deviation | Effect on findings |
|---|---|---|
| 1 | H2.1 (paid vs. unpaid under-5 care) adjudicated from published NSECE snapshots rather than SIPP/NSECE microdata | Precision reduced by an unknown margin; direction unaffected; a full microdata pass remains queued |
| 2, 5, 6 | Scorecard method initially run at reduced independence: same-analyst shuffled sample re-score; 23 dimensions consolidated to 14; sample covered 3 of 12 architectures | #2 and #6 discharged 2026-08-07 by a structurally blinded full-matrix re-score (deviation #10); #5 (14 dimensions) remains |
| 3 | §16 sequencing drafted (v0) before the formal §14 scorecard had run | Sequencing re-checked against the scorecard output this pass; no material discrepancy found, but the ordering is logged as provisional pending formalization |
| 4 | The FTE workforce-requirement model uses representative national licensing ratios and coverage scenarios rather than state-by-state ratio tables | National totals shown robust to ±1 ratio point (sensitivity tested); state-level detail remains queued |
| 7, 8 | Anchor table rows closed incrementally across passes; by pass 3, rows 6, 10, 14, and 15 closed, but rows 4, 7, 9, 11, 16, 17, 18, and 20 remain blank | Two-source rule enforced throughout — blank rows are unfilled, not guessed; none is treated as load-bearing for current conclusions |
| 9 | An academic-literature replication check for the ARPA price-suppression finding was attempted and came back empty — an OpenAlex full-corpus search returned 4 items, none evaluative | Treated as the literature not yet existing, not a search failure; the ws09 conclusion holds at single-source-family confidence, and a Known-Unknowns Register entry was added |
| 10 | Independent structurally-blinded re-score of the full 12×14 matrix (batch API; evidence inlined; scorecard unreachable) | 47 cells corrected, 25 judgment splits logged; stable set is public option / Head Start scaled / K–12; supply-first and caregiver-choice flip the fourth seat; demand-side allowance and Tri-Share rank last under every weighting |
| 11 | Verification Protocol (Phase 1, independent fact-check) artifacts landed from a salvaged branch — extraction ledgers, verification-log.md, steelman suite, source atlas; stale site-chrome edits from that branch were discarded post-#66 | Scorecard evidence-tag claim corrected in the record; full 1,029-row claim ledger with verdicts now committed; public-page fact corrections identified by the pass were logged as still owed (discharged by deviation #12) |
| 12 | Verification Protocol Phase 1 CORRECTED/OVERSTATED findings applied to the whitepaper, this digest, and the scorecard-scales mirror page — Byrd-rule Part 4 rewrite, DOD/Australia/Quebec/Georgia-Florida precedent corrections, workforce and cost figures, Bright Horizons/Tri-Share figures, rural-desert framing, CCDF rule status, and the architecture count (11 + cash comparator, not 12) | Site text now matches verification-log.md's adjudicated verdicts. STALE and UNVERIFIABLE items (population/wage vintages, chain occupancy, CBO 57630, the 220,000-provider figure, DC PEF's FY label, the Treasury attribution, and Germany's REFER-UP courts finding) remain open by design — no replacement figure was fabricated for any of them |
| 13 | Verification Protocol Phase 2 (steelman, built 2026-08-04) ran in the same session as Phase 1's fact-check, deviating from S1's requirement of session independence | Mitigated by building each steelman target in a fresh agent context blind to the checker's reasoning (S2); logged here, not only in steelman-log.md's own preamble, so it can be discounted accordingly. The structural findings below are independently checkable against the committed record regardless |
| 14 | Verification Protocol Phase 2 (steelman) findings landed on the whitepaper and this digest — Part 2's "Binding constraint" stamp qualified, Part 6's "three designs survive every ranking" qualified, and the honesty box extended with the Head Start Impact Study, child_dev_first, and Quebec gaps | Site now discloses where the filing's own instrument and scorecard structure foreclose questions it presents as tested. Corrected 2026-08-11: this entry originally said a3 (Head Start scaled) leads under every weighting and sensitivity tested — a8 (public option) leads all four weightings on the committed board |
Red team
Five attacks on the workforce, market-response, and scorecard conclusions, each answered and, where it landed, absorbed into the record rather than defended against.
Attack 1 — "workforce binds" may be an artifact of studying underpaying programs
Partially lands. Every failing case cited (Canada, Quebec, Germany, UK) set wages low before discovering shortage — a US design paying parity from day one has no clean precedent among the failures. DC's experience supports the attack on retention; but the FTE model's required hiring flow (~798k/yr at ramp) is a labor-market-scale movement no wage level has been shown to produce on schedule. Amended: workforce binds at any historically attempted wage; at parity wages the binding constraint becomes flow and pipeline, which is untested at scale.
Attack 2 — the ARPA price finding is a pandemic artifact from an interested source
Lands on sourcing, not direction. The White House's own Council of Economic Advisers analyzed its own program during pandemic-distorted price dynamics, and the two-source rule isn't cleanly met — the corroborating advocacy source shares the same conclusion but not independent methodology. The post-cliff access deterioration is harder to attribute to pandemic dynamics in 2024, so the direction survives. Amended: downgraded from supported to supported-moderate, single-source-family, pending independent academic replication.
Attack 3 — chain occupancy at 71% shows demand softness, not staffing limits
Lands as segmentation. KinderCare's enrollment fell while prices rose — in chain-served markets, at current prices, demand is genuinely the margin. Both are true in different geographies: deserts (rural, 70%+) are supply-starved; chain-served metros show price-rationed demand with slack. This strengthens the case against a uniform national instrument and for a geography-differentiated design — the inquiry should stop implying universal physical shortage.
Attack 4 — supply-first is politically fatal, and the scorecard buries it
Partially lands. Voters see nothing for years under a pure supply-first sequence, and the durability score for that architecture is low. But Georgia and Oklahoma delivered visible universal seats fast, at modest quality, specifically where school capacity already existed — a demand-first success at the 4-year-old tier, not a rebuttal of supply-first logic at 0–2 where capacity doesn't exist. The sequence is restated as "visible-first where capacity exists, supply-first where it doesn't."
Attack 5 — the stable top-4 reflects correlated scoring, not evidence
Partially discharged by the 2026-08-07 blind re-score. The attack was right that a same-analyst sample could not settle the board: the independent pass moved 47 cells, broke the flip-free top-four, and replaced it with three stable designs plus a flipping fourth seat. What survives is still supply-side / direct-provision / holdout-proof leaning — so correlated priors remain a live concern — but the published stability claim is no longer an unchecked internal consistency result. Per M10, one re-score is "twice checked," not final.
Verification Protocol, Phase 2 — steelman
Where the red team above attacks the filing against itself, Phase 2 of the Verification Protocol attacks it against the strongest opposing case the evidence supports, built with equal effort. Three targets, each derived from the filing's own text (S2), each built by an independently-briefed pass that saw the corrected filing but not the checker's reasoning about where it was weak — a deviation from the protocol's session-independence requirement, logged as deviation #13, run at the user's instruction. Full detail, including the surviving objections and what defeats each steelman in its own evidence, is in the Phase 2 Steelman Log.
Target A — cash and employer instruments should beat supply-side provision
Partially survives. Three of the fourteen scorecard dimensions anchor their scale on the instrument category itself — a cash design scores near the floor by definition, before evidence. Dropping those three dimensions changes the top four under all four weightings: supply-first (a6) exits and caregiver-choice (a11) enters under three, and under equity-first a11 leads outright while the cash comparator takes the fourth seat. What doesn't move: a8 (public option) and a3 (Head Start scaled) hold top-four places under every weighting in both specifications — though a3 leads none of them, a8 leads all four on the committed board. And a real RCT of unconditional cash (Baby's First Years) found no significant effect on maternal work or children's time in care — the steelman's own strongest evidence defeats its labor-supply case. The Head Start Impact Study, commissioned by name in this filing's own protocol, is absent from the record; the Baker–Gruber–Milligan Quebec finding was logged only as a citation count.
Target B — the binding constraint is regulatory (ratios, credentials), not compensation
Partially survives, on the load-bearing word. "Binding" requires that relaxing the constraint alone helps and relaxing others wouldn't — the filing's own definition. Loosening ratios to what Florida's own statute already permits, no pay rise at all, cuts required hiring to ~501k/year, matching what halving turnover saves. The word fails its own test. What doesn't move: the childcare wage relative to its own state's median is uncorrelated with ratio stringency (r = −0.089), so regulation isn't shown to have caused the low pay, and caregiver training remains the best causal defense of credentials found (an 18% cut in accidental child deaths). No scorecard dimension prices regulatory burden and no architecture alters ratios — the alternative was never scored, which is itself the finding.
Target C — is universal non-parental care for ages 0–2 clearly beneficial at all?
Partially survives — and reaches the filing's own answer by an unexamined route. No scorecard dimension measures a child-development effect; the child_dev_first objective is implemented as coverage counts, scoring higher for enrolling more under-2s — the contested proposition, entered as an axiom rather than tested. The filing's own protocol review named Quebec's child-outcome literature and the Head Start fade-out debate as requiring adversarial presentation; neither was delivered (Tennessee VPK was). The substantive literature is genuinely divided — persistent harms above the 81st income percentile in one design, no effect among high-income families in a real lottery, inverted results in another. What doesn't move: setting the coverage-of-under-2s term to zero, rather than assuming it, leaves the top four unchanged — the recommendation survives agnosticism about the very question its instrument can't ask.
All three targets, built independently with no access to each other, converged on the same structural defect from different directions: the scorecard's dimensions encode the filing's conclusions rather than testing them. What survives all of it: a8 (public option) and a3 (Head Start scaled to universal) hold top-four places under every weighting and every sensitivity tested.