GUBMENTPlain talk · policy frontier
Filings / Childcare / Sources / Research Inquiry: Universal Childcare,
GBMT-1 · Research record · No. 1

Research Inquiry: Universal Childcare, Nationwide (United States, Ages 0–12)

childcare/docs/research-inquiry.md
This is a working research document from the childcare filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 1, not a summary of it.

Status: Protocol — not yet executed Scope: United States, federal + 50 states + DC + territories; children birth through age 12 (infant/toddler, preschool, and school-age out-of-school-time care) Lens: Comprehensive — cost, financing, workforce, physical supply, regulation, delivery architecture, demand behavior, market response, private capital and the employer channel, the organized stakeholder field, political economy, legislative and legal mechanics, candidate implementation architectures, sequencing Artifact this produces: A rollout feasibility assessment that names the binding constraints in priority order and evaluates candidate designs against them


0. How to use this document

Each workstream (§2–§16) is written to be executed independently. Every one specifies: the question, why it is load-bearing, the sub-questions, the hypotheses to test, named sources, method, and what "done" looks like.

Framing discipline. This inquiry does not assume universal childcare is desirable, nor that it is achievable. It asks what it would take, and reports honestly if the answer is "more than any plausible political configuration will supply." A protocol that can only conclude "yes, and here's how" is not research.

0.1 Evidence standard

0.2 Search method

The source lists in this protocol are seeds, not the corpus — they are a convenience sample of what the author could recall, which is availability bias operating at the layer where it is least visible downstream. Each workstream therefore begins with a documented search: databases (EconLit, NBER working papers, Google Scholar / OpenAlex / Semantic Scholar, ERIC for the education literature), search strings, date ranges, and inclusion/exclusion criteria — PRISMA-lite. The empirical economics literature in particular (the childcare-market microeconomics canon — Blau, Herbst — and the quasi-experimental universal-program studies — Havnes–Mogstad on Norway, Cornelissen et al. on Germany, Fort–Ichino–Zanella on Bologna, Baker–Gruber–Milligan on Quebec, Gormley's Tulsa studies) must be searched, not recalled; it is the strongest causal evidence in the field and the most underrepresented in policy-shop syntheses.

0.3 Hypothesis discipline

Every hypothesis below receives, before execution begins, written adjudication criteria: refuted if…, supported if…, indeterminate if…. Adjudicating after seeing the data is the forking-paths failure this section exists to prevent.

Worked example, H4.2 (pipeline caps the ramp): refuted if any US jurisdiction has demonstrably expanded its credentialed workforce faster than pipeline throughput via wage increases alone (recruitment from adjacent sectors and returners rather than new credentials); supported if DC's Pay Equity Fund raised wages substantially without a commensurate headcount increase within three years; indeterminate if no jurisdiction has yet raised wages enough to test it.

Directional-balance warning. The hypotheses in this protocol tilt uniformly skeptical — every one predicts difficulty. That is a monoculture, and it receives the same treatment as advocacy optimism: §18 requires a steelmanned feasibility case (built from the red-state universal pre-K record, Canada's live rollout, the DoD system's existence, and the 1971 bill passing both chambers) with the same search effort and evidentiary standard as the steelmanned opposition case.

0.4 Numeric anchors — quarantined

Figures appearing in this protocol are approximate priors, included to orient the search — not findings. Every one is listed in the Anchor Table (§20) with its stated value and location; the workstream that uses an anchor must verify it and record the verified value and the delta. An anchor that was wrong is a finding about this protocol's priors, and the table makes those visible instead of silently corrected. Expect several to be wrong.

0.5 "Binding," operationalized

The deliverable's central promise — constraints ranked, binding versus merely difficult — requires a definition or the ranking is vibes. A constraint is binding for a design if, with all other constraints at observed values, relaxing it alone would materially raise achievable coverage or ramp rate, and relaxing the others would not. Practical test: if the budget doubled, what would stop output from doubling? That is the binding constraint. Apply per design and per age band; the answer plausibly differs across both.


1. Definitional decisions to settle before fieldwork

These are contested, and the answers change every downstream number. Resolve them explicitly in the deliverable's opening section; do not let them stay ambient.

  1. What does "universal" mean? Four distinct designs travel under the word:

    • Universal eligibility, subsidized sliding scale (everyone qualifies, most pay something)
    • Universal free at point of use (Quebec-style flat nominal fee, or genuinely free)
    • Universal entitlement (a legally enforceable right to a slot, as in Germany's §24 SGB VIII)
    • Universal access (supply exists within reach; affordability handled separately)

    These have different costs, different failure modes, and different legal architectures. The deliverable evaluates at least two.

  2. Universal to whom — child or parent? Is the entitlement conditioned on parental work/school activity? Activity tests are the single largest driver of administrative burden and of eligibility churn, and they convert a child-development program into a labor-market program. Decide and defend.

  3. What counts as childcare? Center-based, family child care homes (FCC), family/friend/neighbor (FFN) care, school-based pre-K, Head Start/Early Head Start, before/after-school programs, summer programming, and care for children with disabilities are governed by different rules, staffed by different labor pools, and funded by different streams. A design that only covers licensed center-based care solves perhaps half the problem and probably not the half that binds hardest.

  4. The 0–12 band is not one market. Treat it as three, because supply economics, workforce, and regulation differ qualitatively:

    • Infant/toddler (0–~3): highest ratios, highest unit cost, thinnest supply, least public infrastructure
    • Preschool (~3–5): most existing public infrastructure (state pre-K, Head Start), most political consensus
    • School-age (~5–12): demand is complementary to school — before/after-school hours plus roughly 75 non-school weekdays a year (summer, breaks, holidays, in-service days; the oft-cited "~180 non-school days" counts weekends and overstates the employment-relevant gap). Frequently omitted from "universal childcare" proposals despite being the binding constraint on full-time employment for parents of school-age children.

    Explicitly reject any framing that treats "universal pre-K" as a synonym for the object of study. The 0–3 and 5–12 segments are where this gets hard.

  5. Hours. Standard-day, extended-day, nontraditional-hours (evening/overnight/weekend), and rotating-shift care. A large share of low-wage employment is non-9-to-5; a program built on a school-day calendar systematically misses the workers who most need it. Quantify the mismatch.

  6. What is the program for? Maternal labor supply, child development, family financial relief, and equity of access rank the candidate architectures differently — Quebec optimized labor supply and is criticized on child outcomes; a child-development objective favors quality-conditioned designs that a labor-supply objective rejects as slow and expensive. This protocol does not pick one. The §14 scorecard reports rankings under three explicit objective weightings (labor-supply-first, child-development-first, equity-first) plus an equal-weight baseline, and rank stability across objectives is itself a first-class finding. Any single recommendation must state which objective it assumes and whose choice that is.


2. Baseline: what exists, and what is the denominator?

Question: How many children are in scope, what care do they use now, who provides it, what does it cost, and what is public spending already buying?

Why load-bearing: Nearly every published cost estimate for universal childcare is sensitive mostly to assumptions about take-up among children currently in unpaid or informal care. Get the baseline wrong and every downstream number is wrong by the same factor.

Sub-questions:

Hypotheses to test:

Sources: Census/ACS and CPS (including the CPS child care supplements), SIPP, NSECE (National Survey of Early Care and Education — the key instrument; check its most recent wave and whether it has been re-fielded), NHES After-School Programs and ECPP surveys, HHS/ACF CCDF administrative and market-rate-survey data, Head Start PIR, NIEER State of Preschool yearbook, Child Care Aware price-of-care reports, state licensing databases, BLS QCEW/CES for the child care industry (NAICS 6244), Afterschool Alliance America After 3PM.

Method: Microdata-first. Build the baseline from public microdata (ACS, SIPP, CPS supplements, NSECE) with a reproducible pipeline — code and data vintages committed alongside the report — rather than from secondary reports' toplines. This converts the baseline from citable to checkable. Where sources disagree, present the range and diagnose the disagreement rather than picking one. Document every definitional incompatibility encountered — that list is itself an output.

Done when: The named output exists: a reconciled baseline table (children × age band × state × current arrangement × household income) with per-cell uncertainty bands, a committed code pipeline that regenerates it, and a documented incompatibility log. The deliverable can then state the denominator for every subsequent cost figure. This workstream also completes the first pass of the Anchor Table (§20) for the rows assigned to it.


3. Cost and financing

Question: What does it cost annually at steady state, what does the transition cost, and what revenue instruments can carry it?

Sub-questions:

Hypotheses to test:

Sources: CBO cost estimates and any scoring of prior vehicles (the Build Back Better childcare/pre-K title, ~$400B/6yr as enacted by the House in Nov 2021 — verify figure and structure); Treasury's 2021 The Economics of Child Care Supply in the United States; the Provider Cost of Quality Calculator (PCQC) and successor cost-estimation models; state-level cost studies from NM, VT, DC, and Multnomah County; Center for American Progress and Bipartisan Policy Center cost models; Committee for a Responsible Federal Budget for the fiscal-hawk critique; the academic childcare-market and subsidy literature located via the §0.2 search (Blau's market microeconomics and Herbst's subsidy work as entry points), which supplies elasticities the policy-shop models assume.

Method: Do not build a novel cost model. Instead, collect 4–6 existing published estimates, normalize them onto a common definitional basis (§1), and decompose the differences. The variance between credible models is more informative than any single point estimate, and it identifies which assumptions actually drive the answer.

Done when: The named output exists: a cost-estimate normalization table — one row per published model, one column per normalized assumption (take-up, wage target, ratios, hours, age bands, copay), with each model's headline restated on the common §1 basis — plus a cost range in P10/P50/P90 form with the top three drivers of variance named, and at least three revenue instruments assessed for yield, incidence, stability, and legal durability, including a recession-year stress case (need is countercyclical; payroll and state revenues are procyclical).


4. Workforce (treat as the primary binding constraint)

Question: Where do several hundred thousand to a million additional qualified caregivers come from, and at what wage?

Why load-bearing: Money buys slots only if slots can be staffed. This sector currently loses workers to retail and warehouse jobs paying more with less liability exposure and no credentialing requirement. The working hypothesis of this inquiry is that workforce, not money, is the binding constraint — the research should try hard to falsify that.

Sub-questions:

Hypotheses to test:

Sources: BLS OEWS (SOC 39-9011 childcare workers, 25-2011 preschool teachers) and QCEW; Center for the Study of Child Care Employment (Berkeley) Early Childhood Workforce Index; NSECE workforce component; state workforce registries; DC Pay Equity Fund evaluations; DoD/military child care system compensation structure (the strongest US example of a system that solved retention by paying comparably — study it closely); NAEYC accreditation standards; state licensing ratio tables.

Done when: The named output exists: required FTEs by year by segment (0–3, 3–5, school-age modeled as its own pool) under three wage scenarios, with an explicit maximum feasible ramp rate per segment, and the wage-cost implication fed back into §3.


5. Physical supply and facilities

Question: Do the buildings exist, and if not, how are they built, in a sector that cannot collateralize a loan?

Sub-questions:

Hypotheses to test:

Sources: CAP childcare desert analyses; Reinvestment Fund, LIIF, and other CDFI childcare facility program documentation; state facility fund evaluations; municipal zoning code review (sample ~15 jurisdictions across density types); NIEER and state pre-K facility reports.


6. Delivery architecture, federalism, and program integrity

Question: Who actually runs this, and through what legal instrument?

Why load-bearing: The United States has no national childcare administrative apparatus. The delivery choice determines coverage uniformity, speed, and whether the program survives a change in federal administration.

Sub-questions:

Hypotheses to test:

Sources: CCDF regulations and state plans; Head Start Act and Performance Standards; Medicaid expansion take-up literature as the federalism analogue; GAO reports on CCDF administration; state pre-K governance structures (NIEER); the relevant Establishment Clause line of cases.


7. Quality regulation and the cost–quality frontier

Question: What quality level is being promised, what does it cost, and what is the evidence it delivers?

Sub-questions:

Hypotheses to test:

Method: For each major evaluation, record design, sample, counterfactual condition, effect size, and duration of follow-up. The counterfactual is the crux: a program compared against parental care produces a different estimate than one compared against existing paid care. Sort the literature by counterfactual before comparing findings. Where the literature is genuinely polarized (Quebec outcomes, Tennessee VPK, Head Start fade-out), present each side's preferred specification and what evidence would resolve the dispute — never silently adopt one side's estimate (§0.1's adversarial-presentation rule).


8. Demand: what families actually want

Question: Would families use it, and does the supply being proposed match the demand that exists?

Why load-bearing: Universal programs with low take-up in target groups fail expensively. A supply-side buildout mismatched to parental preference produces empty licensed slots alongside unmet need.

Sub-questions:

Sources: NSECE household component; NHES ECPP and ASPA; state subsidy take-up administrative data; Head Start PIR enrollment vs. eligibility; Afterschool Alliance; Urban Institute work on subsidy access and burden; expulsion literature (Gilliam et al.); unmediated parent and provider voice — public-comment dockets on CCDF rulemakings, state legislative testimony from parents and operating providers, and OIG/GAO interview-based reports. Everything else in this workstream hears families through surveys and intermediaries; these sources narrow that gap (limitation logged in §17).


9. Market response and price effects

Question: What does the existing childcare market do when large public money arrives?

Why load-bearing: This is the most common way a well-funded program underdelivers, and it is chronically under-analyzed in advocacy documents.

Scope boundary: this workstream analyzes how the market responds to public money. §12 analyzes what the market has built on its own and what could be commercialized. Run them together where sources overlap, but keep the two questions distinct in the write-up.

Sub-questions:

Hypotheses to test:


10. Political economy: coalitions, opposition, and durability

Question: What has killed this before, what coalition could pass it, and what makes it survive the next administration?

Sub-questions:

Method: For the four state cases, go to primary sources — enacted statute, implementing regulations, budget documents, and any independent evaluation — and interview-equivalent material (legislative testimony, agency reports). Report what has gone wrong in each, not only what was promised.


Question: Given the rules that actually govern federal lawmaking, which designs can be enacted, and what do those rules force a design to give up?

Why load-bearing: §10 asks whether the votes exist. This workstream asks a different and more constraining question: whether the design that has the votes can survive the procedural machinery. The central working hypothesis is that the rules select for a worse program — that the procedurally viable path (reconciliation) systematically strips out exactly the state-conditioning and standard-setting provisions that make the program work, and forces sunsets that recreate the funding-cliff failure. If true, this is the most decision-relevant finding in the inquiry, because it means the design and the vehicle cannot be chosen independently.

11.1 Senate procedure

11.2 Scoring and budget rules

11.3 Committee jurisdiction and vehicle design

Map which committees own which pieces, because a design that splits jurisdiction multiplies the veto points:

11.4 Constitutional and federalism constraints

11.5 Administrative law and implementation risk

11.6 Statutory collisions with existing law

These are the provisions most likely to be discovered late and to be expensive:

11.7 State-level legislative hurdles

The state path is not procedurally free either:

Hypotheses to test:

Sources: Congressional Research Service reports on reconciliation, the Byrd rule, and Spending Clause conditions (CRS is the best single source here); Senate Budget Committee Byrd rule precedent compilations; CBO scoring methodology documents and the BBB score; the case law named above; GAO reports on CCDF and Head Start administration; NCSL for state supermajority, TABOR, and initiative rules; the Unified Agenda for realistic rulemaking timelines.

Method: Build a procedural feasibility matrix: candidate designs (§14) on one axis, procedural obstacles on the other, each cell scored as survives / survives-if-modified / fails, with the modification named. This matrix is the join between this workstream and the next, and it is the analytical core of the report.

Done when: For each candidate architecture, the report can state the vehicle, the majority required, the provisions expected to be stripped, the litigation exposure, and the realistic time from enactment to first child served.


12. Private capital, employers, and the commercializable frontier

Question: Which components of a solution can be market-provided rather than publicly provided, what have private actors already built, and has their activity moved a nationwide framework closer or further away?

Why load-bearing: A nationwide framework will not be built on empty ground. A large private market already exists, employers already buy childcare benefits, and substantial private capital has already been deployed. That incumbency is simultaneously an asset (real capacity, operating knowledge, a business constituency) and an obstacle (rate-setting politics, concentration risk, and a class of families and employers whose problem is already solved). This workstream also identifies where public money should not go, because the market handles it adequately.

12.1 The organizing hypothesis: Baumol constrains what can be commercialized

Childcare's cost is roughly two-thirds to four-fifths labor (verify against §3's cost decomposition), and the core service — adult attention to small children — is the textbook Baumol case: it cannot be made more productive without becoming a different and worse service. Ratios are the product, not an inefficiency.

Therefore the working hypothesis is that the commercializable frontier lies in the periphery, not in care delivery itself: matching and search, back-office and compliance administration, substitute staffing, facilities finance, benefits administration, and provider-network operations. Software cannot take meaningful cost out of the classroom. Test this hypothesis directly — it determines which components a public framework should build versus buy, and it explains a great deal of the venture-backed failure record in this sector.

Corollary to test: because the periphery is where margin is available, private capital concentrates there while the expensive, low-margin core — infant care, nontraditional hours, rural, and low-income markets — remains underserved. If confirmed, market activity is complementary to a public framework rather than a substitute for it, and the public role is precisely the residual the market declines.

12.2 Market structure and the incumbent private sector

12.3 Employers: what they have tried and what it reveals

12.4 Private money in the problem space

12.5 Has the market helped or hurt a nationwide framework?

Answer this directly rather than leaving it implied. Assess each of the following as a testable claim:

Arguments that market activity has helped:

Arguments that it has hurt or complicated matters:

New adoption challenges the market has created:

Hypotheses to test:

Sources: SEC filings and investor materials for publicly traded operators (the most reliable data on unit economics available anywhere — earnings calls discuss occupancy, wage pressure, and rate sensitivity candidly); CCDF rulemaking comment dockets (providers describe participation barriers in their own words, on the record — the unmediated-provider-voice source for this workstream); PitchBook/Crunchbase for the venture and PE record; IRS Statistics of Income for §45F claim volume; BLS National Compensation Survey and SHRM benefits surveys for employer benefit prevalence; Bright Horizons and other vendor-published employer research (discount appropriately); US Chamber of Commerce Foundation state childcare-and-employers reports; ReadyNation cost-of-crisis estimates (verify figures and note sponsorship); Michigan and Kentucky program evaluations; foundation 990s and published strategy documents; Early Care and Education Consortium materials for the for-profit trade position.

Done when: The report can state which components a public framework should buy rather than build, which market behaviors a framework must anticipate and design against, and a defended answer to whether market activity to date is net-helpful or net-harmful to a nationwide rollout.


13. Civil society, associations, and coalition guidance

Question: What have the organizations closest to this already worked out, where do they disagree with each other, and what does their published guidance supply that this inquiry would otherwise have to derive from scratch?

Why load-bearing: Two reasons, and the second matters more. First, these organizations have collectively produced decades of technical work — model legislation, standards frameworks, implementation toolkits, cost calculators — and some hold primary data available nowhere else. Second, the disagreements inside the pro-childcare coalition are a better predictor of legislative failure than the opposition is. A bill dies when its own coalition splits over delivery model or credential requirements. §10 maps these organizations as political forces; this workstream reads what they actually wrote.

13.1 Whose material to mine

Organize by role, because the same document means different things depending on who produced it.

Professional and standards bodies

Labor

Parent and community

Program constituencies

Policy and advocacy organizations

Intergovernmental associations — the implementation-side voice

The critique — read directly, not in summary

Coalitions

13.2 What to extract

For each organization: (a) stated policy position, including specifics on delivery model, eligibility, work requirements, credentials, and for-profit participation; (b) any model legislation, implementation toolkit, cost calculator, or standards framework — these are reusable design assets; (c) primary data they hold that exists nowhere else; (d) funding sources (cross-reference §12.4); (e) what they have publicly opposed, which is usually more diagnostic than what they support.

13.3 The central output: an intra-coalition fault-line map

Identify and document every axis on which the pro-childcare coalition is genuinely divided. Expected fault lines, to be confirmed and expanded:

For each fault line: who is on each side, how deep the disagreement runs, whether prior bills foundered on it, and whether any drafting approach has successfully finessed it.

Hypotheses to test:

Method note — separate data from position. These organizations produce two very different things: primary data collected nowhere else (price surveys, workforce surveys, state administrative scans), which is evidence about the world; and advocacy framing, which is evidence about coalition positions. Both are useful, for different purposes. Do not cite advocacy collateral as though it were research, and do not discard primary data because an advocacy organization collected it. Tag every source with which kind it is.

Done when: The fault-line map is complete, reusable design assets (model bills, toolkits, standards frameworks) are catalogued rather than re-derived, and the opposition case exists in a form its own proponents would endorse as accurate.


14. Prospective implementations: candidate architectures

Question: What are the actual, concrete ways this could be built, and how do they score against the constraints found in Workstreams 1–12?

Why load-bearing: "Universal childcare" is not a policy; it is a category. The inquiry is only useful if it terminates in named, specifiable architectures that can be compared. This workstream is where the analysis converts into options.

14.1 Inventory what is already on the table

Before designing anything, establish what has been introduced and what is currently live. My prior knowledge of the current bill landscape is stale and possibly wrong — treat every item below as a search term, not a fact, and pull the current inventory from Congress.gov directly.

Known bill families to locate, verify status, and read (sponsors, structure, CBO score if any, committee action, cosponsor counts and their partisan composition):

Also inventory the administrative-action space: what could be done without legislation at all — CCDF rulemaking on payment practices and rate-setting methodology, waiver authority, Head Start Performance Standards revisions, procurement and federal-employee childcare expansion. This is the low-ceiling but fast path, and it deserves an honest assessment rather than dismissal.

14.2 Evaluate candidate architectures on a common scorecard

Specify each architecture concretely — funding mechanism, administering entity, eligibility rule, payment method, standards source, and treatment of each of the three age segments — then score all of them on the same criteria.

Architectures to evaluate (at minimum):

  1. CCDF supercharge. Expand and restructure the existing block grant; convert to capped mandatory funding; reform rate-setting to cost-estimation. Least new machinery, fastest to stand up, inherits all existing state variation and the non-participation problem.
  2. Medicaid-model entitlement. Open-ended federal-state match, state plan under federal standards, individual entitlement. Uniform benefit, strong durability, maximum federalism exposure post-NFIB, and heavy Byrd exposure on the standards.
  3. Head Start model, scaled to universal. Federal-to-local grantee network bypassing states entirely. Solves the non-participating-state problem outright; requires building a grantee network of unprecedented size; existing Head Start constituency is a potential ally or a serious obstacle depending on handling.
  4. K–12 extension. Public education extended downward to 3 and outward to before/after-school and summer, run by districts on existing governance, facilities, and funding formulas. Cheapest capacity available and strong durability; does not reach 0–3 at all; collides with the mixed-delivery provider base and with union/labor-agreement structures.
  5. Demand-side allowance or advanceable refundable credit. Money to families, no supply-side machinery. Procedurally the easiest by a wide margin — single committee, clean reconciliation fit. Highest pass-through and price-inflation risk per §9; addresses none of the binding constraints identified in §4 and §5.
  6. Supply-first build-out. Fund workforce compensation, credentialing pipeline, and facilities for 3–5 years before opening universal demand-side eligibility. Directly targets the constraint the inquiry expects to bind; politically the hardest to sell because voters see no benefit for years.
  7. Social-insurance / payroll-tax fund. A dedicated childcare fund on the paid-leave model, with Vermont Act 76 as the state analogue. Dedicated revenue improves durability and reduces appropriations exposure; regressive-incidence critique must be answered.
  8. Public option. Direct public provision alongside the existing market, DoD-style, plausibly anchored in federal facilities and federal employment. Solves supply and workforce directly by paying properly; politically the hardest; strongest quality control.
  9. Hybrid: federal fallback ladder. Federal-state match with strong incentives, plus direct-federal provision in non-participating states — the ACA exchange structure applied to childcare. Likely the pragmatic frontier; evaluate whether it actually clears §11.4.
  10. Blended employer/state/family cost-sharing at scale. The Tri-Share model (§12.3) federalized — three-way cost splitting with a federal match replacing or supplementing the state share. Recruits employers as paying partners rather than lobbyists, and has bipartisan surface appeal; inherits the per-employer administrative cost problem (H12.4) and, if H12.3 holds, may entrench the employer-linked structure that suppresses demand for universality. Evaluate on both counts.
  11. Caregiver-choice allowance. Public money payable to the caregiver the family chooses — including parents and kin — on the home-care-allowance model (Finland is the canonical case; several US states already permit relative-caregiver payment at CCDF's margins). The design a family-values constituency can support, and the only architecture that reaches the families §8 shows prefer home care for infants. The Nordic evidence cuts hard the other way on maternal employment — home-care allowances measurably depress it. Score with that trade named, not hidden. Note: §12.4's funder-bias check predicted this design would be systematically under-analyzed; its original omission from this very list is the confirmation.

Benchmark (not an architecture): the cash-equivalent comparator. Score an equivalently funded, non-earmarked child benefit (CTC-style cash) on the same dimensions as the eleven architectures. It is not a childcare program and will score poorly on supply dimensions — that is the point: it is the identification strategy for what care-specific provision adds per dollar, and it is the comparison opponents will make in every hearing. A report that has not made it first is unarmed.

Design-for-learning requirement. Each architecture specification includes its embedded evaluation design: staged implementation as an evaluation instrument (lottery allocation where oversubscribed; staggered geography enabling difference-in-differences), a pre-registered outcome dashboard tied to §16's leading indicators, and named decision points where evidence redirects the ramp. The rollout is the research. An architecture that cannot be evaluated while operating is a worse architecture, and scores accordingly.

A note on the public/private division of labor. Per §12.1, score each architecture not only on what government provides but on what it buys rather than builds — compliance tooling, provider-network operations, substitute staffing, facilities finance, benefits administration. An architecture that assumes the public sector constructs all of this from scratch is understating both its cost and its timeline; one that assumes a private operator will appear in thin markets is contradicting the §12.4 failure record. State the assumption explicitly for each.

Scoring method (fixed before any scoring begins — §19's structure-before-results rule):

  1. Each dimension gets an anchored ordinal scale — a written description of what a 1, 3, and 5 concretely look like — authored before the first architecture is scored.
  2. Every cell cites the workstream finding it rests on. An uncited cell is a flag, not a score.
  3. No single weighted ranking. Report rankings under the three §1.6 objective weightings plus an equal-weight baseline, and report rank stability: where the ranking is invariant across objectives versus where it flips is the most decision-relevant output of the exercise.
  4. Score the matrix twice, independently — two analysts, or the same analyst re-scoring from a shuffled architecture order after a gap — and reconcile disagreements in writing. The disagreement log is a deliverable.

Scorecard dimensions (score every architecture on every dimension — an architecture that is not scored on a dimension has not been evaluated):

Dimension Source
Which binding constraint does it actually relieve? §4, §5
Steady-state and ramp cost §3
Coverage: what share of children in each of the three age bands §1, §2
Coverage uniformity across states §6, §11.4
Time from enactment to first child served §11.5
Maximum feasible ramp rate §4, §5
Legislative vehicle and majority required §11.1, §11.3
Byrd-rule survival of core provisions §11.1
Constitutional and litigation exposure §11.4, §11.6
Durability across a change of administration §10, §11.2
Price pass-through and market-distortion risk §9
Provider-base effects: who exits, who consolidates §9, §12.2
Incumbent participation risk: will chains and PE-owned operators accept the rate? §12.5
Public/private division of labor: what is bought rather than built §12.1
Effect on the employer-benefit constituency (does it entrench or dissolve it?) §12.3
Which intra-coalition fault lines it triggers §13.3
Administrative burden on small and home-based providers §11.6, §12.2
Handling of nontraditional hours and rural supply §5, §8
Evidence base supporting the expected outcomes §7
Distributional incidence: who gains, by income, race/ethnicity, geography, and work schedule §2, §8, §12.3
Integrity and enforcement risk profile: fraud exposure and clawback-harm exposure §6
Resilience in a recession: countercyclical need vs. procyclical revenue §3, §16
Embedded evaluation design: can it be evaluated while operating? §16

Hypotheses to test:

Done when: The scorecard is complete, at least two architectures are specified in enough detail to be drafted into legislative text, and the report can state which architecture is recommended for each age band and why — including the case against the recommendation.


15. International and domestic precedents

Question: What can be learned from systems that already did this, adjusted for transferability?

Cases (each gets: design, cost as % GDP, coverage achieved, ramp time, workforce solution, evaluated outcomes, known problems):

Method: Use a common comparison template so cases are actually comparable. For every case, include a transferability assessment naming the structural preconditions the US does not share. A case study without that section is decoration.


16. Sequencing and transition

Question: Given that everything above cannot happen at once, what is the correct order, and what is the ramp rate?

Why load-bearing: This is where the inquiry converts into something actionable. The constraints identified in §4 (workforce pipeline) and §5 (facilities) impose a maximum feasible ramp regardless of appropriation size, and §11 imposes a second ceiling — time from enactment through rulemaking and state plan approval to first child served. Money exceeding that rate produces price inflation, not capacity — which then discredits the program.

Sub-questions:

Deliverable for this workstream: A sequencing recommendation with year-by-year milestones for years 1–10, explicit dependency ordering, and the leading indicators that would tell you the plan is failing early enough to correct it.


17. Known-Unknowns Register

Maintain throughout execution. Each entry: the question, why it cannot currently be answered, what would be required to answer it, and how much the overall assessment depends on it.

Seed entries (expected, to be confirmed or removed in §2):


18. Deliverable specification

Primary output: A rollout feasibility assessment, structured as:

  1. Executive summary — the binding constraints in priority order, and the single most important thing that must be true for this to work
  2. Definitions and design space (§1) — the two or three designs evaluated, stated precisely
  3. Baseline (§2) — the denominator, with uncertainty bands
  4. Constraint analysis (§4, §5, §6) — workforce, facilities, delivery architecture, each with a derived maximum feasible ramp
  5. Cost and financing (§3, §9) — range, drivers of variance, revenue instruments with incidence
  6. What we would be buying (§7, §8) — quality evidence and demand realism, with counterfactuals stated
  7. Precedents (§15) — comparison table plus transferability assessments
  8. The private market and the employer channel (§12) — what is already built, what can be commercialized, whether market activity is net-helpful, and the rate-participation and crowd-out risks a framework must design against
  9. Stakeholder landscape (§10, §13) — coalition and opposition map, the intra-coalition fault-line map, catalogued reusable design assets, and a steelmanned opposition case
  10. Legislative and legal mechanics (§11) — the procedural feasibility matrix, the vehicle-versus-design tradeoff, and time-to-first-child-served
  11. Candidate architectures (§14) — the full scorecard under multi-objective weightings with rank-stability analysis, at least two designs specified to draftable detail, the cash-equivalent comparator, and a recommendation per age band with the objective it assumes stated
  12. Sequencing recommendation (§16) — 10-year milestones with dependency ordering, leading indicators, and recession stress tests, covering legislative as well as operational sequence
  13. Known-unknowns register (§17) and the completed Anchor Table (§20)
  14. What would change the conclusion — the specific findings that would flip the recommendation

Format: Written report. Every number sourced and dated. A separate one-page summary of binding constraints, suitable for someone who reads nothing else.

Required properties:


19. Execution notes

Phase-gate. Do not commit to full execution up front. Run a Phase 0 first: verify the §20 Anchor Table's highest-leverage rows, build the §2 baseline skeleton, and scan the §15 additions (red-state universal pre-K, Canada's rollout). Phase 0's job is to fire the kill conditions early or re-rank the workstreams while it is still cheap — several of this protocol's priors could be reshaped by a few days of verification, and it is cheaper to find that out in days than months.

Suggested order. §2 (baseline) first — everything else depends on it. Then §4 (workforce) and §15 (precedents) in parallel, since the working hypothesis is that workforce binds and the state/international cases are the strongest available evidence on whether that is true. Then §3, §5, §6 in parallel. Then §7, §8, §9, with §12 (private market and employers) alongside them — §12 is a direct extension of §9's market analysis and shares most of its sources. §13 (civil society) can run at any point after §10 and is highly parallelizable. Then §11 (legislative mechanics), which needs §6's delivery-architecture options in hand before it can test them against the rules. Then §14 (candidate architectures), the synthesis point where every prior workstream feeds the scorecard. Then §16 (sequencing) last.

§11 and §14 are coupled. Do not execute them separately: the procedural feasibility matrix in §11 and the architecture scorecard in §14 share axes, and running them as one exercise avoids scoring designs that the rules have already eliminated.

§12 and §13 have a natural division of labor and should not be merged. §12 asks what the market and employers have done; §13 asks what the organized stakeholder field has written. Both touch the same organizations from opposite directions — the Chamber Foundation and ReadyNation appear in each — so cross-check rather than duplicate.

Versioning and deviations. The protocol is under version control; the version at execution start is the freeze point. Every mid-course deviation — dropped sub-question, added source type, reinterpreted hypothesis — is logged with a reason in a deviations file committed beside the report. Silent drift toward whatever the data made easy is the failure mode this prevents.

Structure before results. The report's section structure, decision criteria, scorecard scales, and objective weightings (§1.6, §14.2) are fixed before synthesis begins, so the frame cannot be fitted to the findings.

Red-team pass. Before finalization, one reviewer's sole brief is to attack the draft's conclusions using §13's opposition sources and the strongest contrary academic findings, in writing; the final report answers or absorbs each attack, also in writing.

External review before execution. One outside read each from a childcare-policy academic, a budget-process practitioner, and a current or former state subsidy administrator. Their objections will not overlap with any self-review's.

Run §12.4's funder-bias check against this inquiry itself at the end of execution: which options received the least analytical attention, and does the pattern track the funding structure of the sources used? It has caught one blind spot already — the caregiver-choice allowance's original omission from §14.2.

Kill conditions. If §2 shows the baseline data cannot support the analysis, stop and report that as the finding rather than proceeding on invented numbers. If §4 shows the workforce ramp cannot reach the scale required under any wage assumption, that dominates the rest of the analysis and should be reported immediately rather than at the end. If §11 shows that every architecture relieving a binding constraint is procedurally dead, report that as the headline — it reframes the whole question from "what should we build" to "what has to change before anything can be built."

Effort shape. Roughly half of total effort on §2, §4, §11, and §14 — baseline, workforce, legislative mechanics, and candidate architectures carry the most decision-relevant information per hour. §12 (private capital and employers) has the highest ratio of unread primary material to conclusions currently in circulation; operator SEC filings and earnings calls in particular are underused and unusually candid about unit economics. §10 (political economy) is the most likely to expand without adding decision-relevant content; time-box it, and push anything that is really a rules question into §11 and anything that is really a published-position question into §13.

Currency warning for §14.1. The bill inventory changes every Congress and the sponsor names, bill numbers, and statuses in this protocol are unverified recollections. Pull the current picture from Congress.gov before doing anything else in that workstream, and treat the named bills as search terms only. The same warning applies to the company, chain, ownership, and funder rosters in §12, the organization list in §13, and the precedent-program details in §15 — ownership, coalition membership, and program parameters all turn over quickly.


20. Anchor Table

Every numeric prior stated in this protocol, quarantined per §0.4. The workstream assigned to an anchor verifies it and completes the row; deltas are findings about this protocol's priors. Phase 0 (§19) makes a first pass over the highest-leverage rows (marked ★).

# Anchor (as stated) Where used Verify in Verified value & source Delta
1 ★ BBB childcare/pre-K title ≈ $400B / 6 yr as House-passed (Nov 2021) §3 §3 CBO: +$381.5B deficit 2022–2031 for the childcare+preschool provisions; $400B was the bill's line item. CBO assumed substantial state non-participation (CBO 57630) Anchor conflated line item with score; non-participation assumption is a §11.4 finding
2 ★ ARPA stabilization ≈ $39B, expired 2023-09-30 §9, §10 §9 $39B = $24B stabilization + $15B CCDF supplemental; stabilization expired 2023-09-30, supplemental wound down through Sept 2024 (ACF) Only the $24B expired on that date; cliff was two-stage
3 ★ Childcare cost is ⅔–⅘ labor §12.1 §3 Treasury 2021: wages ≥50–60% of expenses on US averages, higher for infant care; center personnel commonly cited at 70–80% (Treasury) Widen range to 50–80%; varies by age band and setting
4 ≈75 employment-relevant non-school weekdays per year §1 §2
5 ★ Additional workforce needed: several hundred thousand to ~1M §4, §9 §4 Baseline ≈1.05M industry jobs; FTE model (pass 2): 2.8M FTEs required central (range 1.9–3.9M), net new ≈1.76M — see baseline/scripts/fte_model.py Prior understated requirement by ~2× — the anchor-quarantine table doing its job
6 Tax-increase supermajority requirements in ≈12 states §11.7 §11 17 states require legislative supermajorities for tax increases (FGA; NCSL primary confirmation queued) Prior understated by ~5 states — state-revenue path harder than assumed
7 Ballot initiative available in ≈half the states §11.7 §11
8 CCDBG authorization lapsed after the 2014 reauthorization period §11.2 §11 Verified: 2014 reauthorization (P.L. 113-186) authorized FY2015–FY2020; expired; funded by appropriations since (CRS R47312) Confirmed
9 FRA 2023 discretionary caps covered FY24–25 §11.2 §11
10 Dependent care FSA cap nominally near-frozen for decades (brief ARPA-era increase) §12.3 §12 Verified and superseded: $5,000 set by Tax Reform Act of 1986, never indexed — raised to $7,500 effective 2026 (OBBBA, signed 2025-07-04) (Newfront, EBC) Anchor was true and just became stale — employer/tax channel now expanding on two tracks (§45F + DCFSA), H12.3 live
11 §45F persistently under-claimed relative to authorization §12.3 §12
12 ★ Georgia universal pre-K mid-1990s; Oklahoma late 1990s; Florida mid-2000s §15 §15 GA 1995 (first state; lottery-funded; ~55% of 4s; 60% of classrooms private — DECAL); OK 1998 (school-funding-formula route; >70% of 4s — the nation's highest); FL voter-approved by ballot, program operating 2005 (New America) Confirmed; pin FL amendment vs. launch dates in full §15 pass
13 ★ Canada CWELCC: 2021 agreements, $10/day target §15 §15 Targets missed: 194k of 284k new spaces (Sept 2025); Ontario fees ≈$19/day, not $10; deadline extended to Dec 2026; Ontario short up to 10k ECEs; 57% of new spaces for-profit, Ontario pushing to lift the cap (Globe, CCPA) Live confirmation: workforce binds the ramp (§4's hypothesis) and the §12.5 for-profit fight is real
14 Netherlands scandal: thousands of families falsely accused; government fell (2021) §6, §15 §15 Verified, worse than stated: ~26,000 families (estimates to 35,000) wrongly ordered to repay, ethnic profiling documented, 2,000+ children removed into custody; government resigned Jan 2021 (Wikipedia/NL Times) Clawback-harm case even stronger than the prior
15 NYC pre-K: tens of thousands of seats in ≈2 years §15 §15 Verified: funded expansion 20k→53k full-day seats (2014); >60k enrolled within two years (NYC ODA, TCF) Confirmed — the domestic rapid-ramp existence proof
16 Background-check backlogs of months post-2014 CCDBG in some states §4 §4
17 Michigan Tri-Share: ≈equal three-way split; Kentucky employer match §12.3 §12
18 Quebec reform 1997; Germany slot entitlement from age 1 in 2013 §15 §15
19 1971 CCDA passed both chambers, vetoed; CCDBG enacted 1990 §10 §10 Verified: vetoed 1971-12-10; veto message (drafted by Buchanan): "communal approaches to child rearing over against the family-centered approach"; "family-weakening implications" (APP primary text) Confirmed, primary source
20 Head Start since 1965; Lanham centers 1943–46 §15 §15
← All Childcare research documents Sources digest Read the whitepaper