Target C: universal publicly-funded non-parental care for ages 0–2 is not clearly beneficial and may be harmful
Run per method/verification-protocol.md §S1–S5. Built from the filing's own text (S2) and then from primary sources fetched in this session (S3). Every figure below was read inside a document I fetched myself unless explicitly marked otherwise.
1. S2 — The tilt, written before building anything
The tilt statement.
GBMT-1 leans pro-universal-coverage on the goal and skeptical on the means. Every one of its ~30 hypotheses, its five red-team attacks, its adjudication criteria, and its own steelman are about whether universal childcare for ages 0–12 can be built — never whether it should be. Universal coverage of all 51.2 million children 0–12, including the 11.1 million under three, enters as the definition of the problem ("A feasibility assessment of nationwide childcare for the 51.2 million American children ages 0–12"), and the filing's conclusion is a construction schedule: "feasibility is a design and sequencing problem — pay the workforce first, build supply second, extend benefits third." The position it never entertains is that a universal offer at 0–2 may be the wrong shape of policy on child-outcome grounds, independent of whether it can be staffed.
Four pieces of the filing's own record establish this, and they are load-bearing for §4 below.
The filing's own steelman reveals what it treated as the unfashionable side.
steelman-feasibility.mdis titled "The Steelmanned Feasibility Case" and opens: "The strongest honest argument that nationwide universal childcare is more achievable than this inquiry's skeptical priors assume." The filing's method (M7) requires steelmanning both directions; the direction it chose to steelman was optimism about the build. The symmetric case it did not build is pessimism about the goal. Its own closing line — "That is the difference between 'hard' and 'infeasible,' and the evidence supports 'hard'" — is an answer to a feasibility question only.No pre-registered hypothesis can return "the goal is wrong."
adjudication-criteria.mdfixes criteria for seven hypotheses: H2.1 (share of under-5s in paid care), H4.1 (wage parity as share of cost), H4.2 (pipeline caps the ramp), H9.1 (price absorption), H11.1 (Byrd rule strips universality-plus-quality), H12.3 (employer crowd-out), H14.2 (segmented beats uniform). Every one is about cost, supply, procedure, or architecture selection among ways of achieving universal coverage. None has a "supported" branch that means "do not do this." The nearest thing in the wider record, H7.2 ("scaled programs produce smaller effects than the demonstration programs used to justify them"), is about the magnitude of a presumed benefit, not its sign, and it is not in the pre-registered criteria file at all.The scorecard has no child-outcome dimension — and its "child development" objective is defined as coverage.
scorecard/scales.mdlists 14 anchored dimensions: relief_workforce, relief_supply, coverage_02, coverage_34, coverage_512, procedural, federalism, durability, passthrough_risk, participation_risk, preference_fit, integrity_risk, distributional, evaluability. Not one is "effect on child development." Andscorecard/rank.pyimplements thechild_dev_firstobjective as:"child_dev_first": {**{d: 1 for d in dims}, "relief_workforce": 3, "coverage_34": 2, "preference_fit": 2, "evaluability": 2, "coverage_02": 2},Every up-weighted term is an input or a coverage count. Under this objective an architecture scores better for child development precisely by covering more under-2s (
coverage_02scale: 5 = "Universal 0–2 offer"). The proposition this steelman contests is not merely unexamined by the scorecard — it is encoded as an axiom of the scorecard's own child-development objective. The filing'sprotocol-review.md§A3 warned exactly this would happen ("the scorecard has no objective attached, which means the 'recommendation per age band' that §14 promises cannot be derived from the scorecard — it will be smuggled in"), and the fix adopted (three objective weightings) answered the symptom while leaving the dimension set unchanged.The literature was retrieved and then dropped.
docs/phase0-findings.md:38records, as a success: "OpenAlex's free API (no auth) returned the §0.2 canon on the first query — Baker–Gruber–Milligan (NBER w11832; 271 citations), Cornelissen et al. (JPE 2018), and the BGM long-run follow-up (AEJ:Policy 2019)." Those three papers are the core of the causal literature on whether universal childcare helps or hurts children. None appears in any workstream file, the anchor table, the scorecard, or either public page.ledger-record.md:R210carries the entry "Quebec child outcomes are contested; literature is genuinely polarized" — sourced fromprotocol-review.md, with the source column reading "none stated."protocol-review.md§A5 predicted the omission ("the quasi-experimental universal-program literature beyond Quebec — Havnes–Mogstad on Norway, Cornelissen et al. on Germany, Fort–Ichino–Zanella's negative findings for affluent children in Bologna") and recommendation 8 required "adversarial presentation for polarized literatures … (Quebec child outcomes, Tennessee VPK, Head Start fade-out)." The filing ran that adversarial presentation for Tennessee VPK vs. Boston inws07-quality.md. It did not run it for Quebec.ws15-precedents.mduses Quebec for exactly two facts — a 30,688-child waitlist and a staffing attribution — under the heading "the queue is permanent, and staffing is why."This is not an availability failure. The record shows the retrieval worked, the omission was predicted in advance by the filing's own reviewer, and the material was dropped anyway. That is a tilt, not an accident.
One point of fairness owed to the filing before proceeding. ws07-quality.md came within one step of this question and turned away from it deliberately, not ignorantly. It reports Tennessee VPK's RCT negative effects and concludes: "quality parameters (curriculum, coaching, workforce stability) are not garnish — they plausibly determine the sign of child effects." The filing therefore knows the sign can be negative. It routed that finding into a design parameter ("this raises the stakes on §11's finding that quality provisions are the most Byrd-vulnerable part of any bill") rather than into a goal question. The framing did the excluding, and it did it with the relevant fact in hand.
2. The steelman
The case has four legs. It is not "childcare is bad for children." It is narrower, and that is what makes it survivable:
At ages 0–2 specifically, the sign of the effect of non-parental centre care on child development is a function of what the child would otherwise have been doing. A universal entitlement is the single instrument that maximises the number of children moved out of the counterfactual where the return is highest — a resourced home — and into the setting where the measured return for those children is zero or negative. The well-identified literature does not support universality at 0–2 on child-development grounds; it supports targeting. The filing's framing cannot see this because its objective function treats 0–2 coverage as a child-development good by construction.
Leg 1 — Quebec: the longest-running universal programme in North America, and its outcome literature
The filing cites Quebec once, for a waitlist. Here is what the economics literature on Quebec's child and family outcomes actually found.
1.1 Baker, Gruber & Milligan — contemporaneous effects
Universal Childcare, Maternal Labor Supply and Family Well-Being, NBER WP 11832 (Dec 2005), published JPE 116(4):709–745 (2008). [Fetched: nber.org/system/files/working_papers/w11832/w11832.pdf]
- Identification: difference-in-differences, Quebec vs. rest of Canada, before/after the 1997 $5/day reform, National Longitudinal Survey of Children and Youth, two-parent families with a preschool child. Placebo group: 6–11-year-olds (largely outside the treatment).
- Findings (abstract, verbatim): "we uncover striking evidence that children are worse off in a variety of behavioral and health dimensions, ranging from aggression to motor-social skills to illness. Our analysis also suggests that the new childcare program led to more hostile, less consistent parenting, worse parental health, and lower-quality parental relationships."
- Specifics I read in the results section: positive and significant effects on the pooled hyperactivity score, on anxiety (both age bands), and on aggressiveness (both age bands); a highly significant negative effect on motor-social skills; a significant −5.3 percentage point effect on the odds of being in excellent health; large significant reductions in the odds of never having infections. PPVT (cognitive, age 4) positive, small, imprecise — "uninformative."
- Magnitude discipline, in the authors' own words: these are ITT estimates on all children. Bounding the treated share at 7.7%–19.5%, motor-social skills fall 8.4%–21.2% relative to the mean for the treated; anxiety rises 62.6%–158.5%. They calibrate against controls: "the CPE program's effect on hyperactivity — even when we take the lower bound estimate — is larger than the effect of mother's high school completion or the boy-girl differential."
- Placebo behaves: for three of six measures available for 6–11-year-olds the estimates are wrong-signed.
- A finding the filing would have wanted: "approximately one third of the newly reported use appears to come from women who previously worked and had informal arrangements" — i.e. a third of the "new" childcare was crowd-out of informal care, and the authors compute that only ~40% of programme cost was recovered by income and payroll taxes on the induced labour supply.
1.2 Baker, Gruber & Milligan — long run
The Long-Run Impacts of a Universal Child Care Program, AEJ: Economic Policy 11(3):1–26 (2019); NBER WP 21571. [Fetched both: published PDF via dspace.mit.edu/bitstream/handle/1721.1/126667/pol.20170603.pdf, WP via NBER]
- Identification: same DiD design extended to cohort exposure; NLSCY at ages 5–9, CCHS/CMHS for teen self-reports, Statistics Canada Uniform Crime Reporting Survey for province × age × year × sex crime rates.
- Non-cognitive deficits persist and in places grow. At ages 5–9: anxiety "just over one quarter of a standard deviation, which is more than twice as large as for 2–3-year-olds"; hyperactivity +13% SD; indirect aggression +19% SD; aggression similar to the younger estimate. Parent report of the child getting along with the teacher significantly worse. In the WP's gender split, aggression effects are "primarily for boys, and the estimate in this case is one third of a standard deviation."
- Health and life satisfaction (ages 12–20): poor-health indicator +7.3% SD; life satisfaction significantly worse; self-reported mental health negative but small and insignificant.
- Crime. In the published version's richest specification: accusations across all categories +353 per 100,000 against a mean of 1,872 — a 19% rise; convictions 22%. Property crime accusations +602 per 100,000 (19% of mean), convictions +342 (≈25% of mean); crimes against persons 16% of mean; drug convictions >23% of mean. Effects larger in absolute terms for boys.
- They handled the obvious confound. Quebec diverts youth offenders differently, and the Youth Criminal Justice Act (in force 1 April 2003) made the rest of Canada more like Quebec. The published paper states it uses crime data starting in 2006 specifically "to stay clear of this impact of the YCJA," and examines both accusations and convictions to bracket the effect of extrajudicial remedies.
- Honest about cognition. "On balance, the results in Table 3 do not provide unambiguous evidence of a persistent negative impact of the Quebec program on cognitive ability." One PISA specification shows math +30% SD. The damage claim is about non-cognitive development, not test scores.
1.3 The replication that tried to break it
Kottelenberg & Lehrer, New Evidence on the Impacts of Access to and Attending Universal Childcare in Canada, NBER WP 18785 (Feb 2013) → Canadian Public Policy 39(2). [Fetched: NBER]
Abstract, verbatim: "we show the robustness of the initial analyses to i) concerns over whether negative outcomes would vanish over time as suppliers gained experience providing child care, ii) concerns regarding multiple testing, and iii) concerns that the original test measured the causal impact of childcare availability and not child care attendance." This is the field's most serious attempt to overturn BGM, and on its own account the core survived. (Its two important qualifications are in §3 below — they cut against me and I state them there.)
1.4 The family-functioning channel, independently reproduced
Haeck, Lebihan, Lefebvre & Merrigan, Universal child care and longer-run effects on parental health and behaviors, UQÀM GRCH WP 15-04 (Nov 2015). [Fetched: grch.esg.uqam.ca]
"We show that the policy increased mothers' depression scores with preschool children as well as scores of inappropriate parenting behavior. The policy increased hostile and aversive parenting and reduced positive interaction and consistent parenting." A separate team, separate specification, same NLSCY, replicating BGM's parenting result. (Their qualification — the effects vanish once the child is in school — is in §3.)
What Leg 1 establishes: the single longest-running universal childcare programme in North America produced a documented, replicated deterioration in child non-cognitive development and parenting quality, with long-run correlates in health, life satisfaction and crime. Whether it nets out negative is contested (§3). That it is contested is not.
Leg 2 — Age at entry and intensity for under-3s: the same result outside Quebec
2.1 High-quality care, affluent families, regression discontinuity — still negative
Fort, Ichino & Zanella, Cognitive and Noncognitive Costs of Day Care at Age 0–2 for Children in Advantaged Families, JPE 128(1):158–205 (2020). [Fetched: University of Chicago Press manuscript preprint, DOI 10.1086/704075, stamped "Copyright The University of Chicago 2019"]
- Identification: RDD on the Bologna Daycare System's admission thresholds. Applicants are ranked within priority group by a household-size-adjusted "Family Affluence Index"; capacity sets a threshold; applicants just below get their preferred programme, just above do not. Outcomes measured by professional psychologists at ages 8–14 (WISC-IV for IQ, BFQ-C for Big Five), 2013–15, on 2001–05 applicants.
- Finding: "an additional month in daycare at age 0–2 reduces IQ by about 0.5%, on average. At the sample mean (116.4), this effect corresponds to 0.6 IQ points (4.7% of the IQ standard deviation) and its magnitude increases with family income." For better-off families, an extra month reduces agreeableness and openness by ~1% and increases neuroticism by ~1%.
- This kills the "it was just low quality" defence. Bologna's asilo nido system is described by the authors as "renowned for its high-quality even outside the country," with an adult-to-child ratio of 1:4 at age 0 and 1:6 at ages 1–2 — better than almost any US licensing standard. The comparison is not bad daycare vs. good home. It is good daycare vs. a counterfactual ("parents, grandparents, and nannies") with an adult-to-child ratio "close to 1."
- And it supplies the mechanism that unifies the whole literature. Their model: "when daycare time increases, child skills decrease in a sufficiently affluent household because of the higher quality of home care … For a less affluent household, instead, the offer of the most preferred program increases both household consumption and child skills, because home care is of a lower quality than daycare."
2.2 US observational evidence on hours, tracked to age 15
Vandell, Belsky, Burchinal, Steinberg & Vandergrift, Do Effects of Early Child Care Extend to Age 15 Years?, Child Development 81(3):737–756 (2010), NICHD SECCYD. [Fetched raw: pmc.ncbi.nlm.nih.gov/articles/PMC2938040/]
- N = 1,364, birth cohort recruited 1991, nonrelative care birth to 4½ years, outcomes at 15.
- Verbatim from the results: "Adolescents who experienced more hours of nonrelative child care across their first 4½ years reported significantly more risk taking (B = .008, d = .09) and greater impulsivity (B = .08, d = .13) at age 15. In addition, children who experienced higher quality care had significantly lower externalizing scores (B = −1.89, d = .09)."
- The load-bearing asymmetry: quality buys achievement and lower externalizing; hours buy risk-taking and impulsivity, and the hours effect survives controlling for quality. The authors note these are "surprisingly similar" in magnitude to the hours→externalizing effect first detected at 4½ (d = .16), i.e. a decade of persistence.
- Stated honestly by the authors, and I repeat it: "a large, nonexperimental field study … which affords estimation of statistical rather than causal effects. When causal language (e.g., effect, influence) is employed in this report, it is for heuristic purposes."
2.3 The US subsidy design itself
Herbst & Tekin, The Impact of Child Care Subsidies on Child Well-Being: Evidence from Geographic Variation in the Distance to Social Service Agencies, NBER WP 16250 (Aug 2010) → Journal of Public Economics. [Fetched: NBER]
- Identification: IV using distance from home to the nearest social service agency administering the CCDF subsidy application, ECLS-K kindergarten cohort.
- Finding, verbatim: "children receiving subsidized care in the year before kindergarten score lower on tests of cognitive ability and reveal more behavior problems throughout kindergarten. However, these negative effects largely disappear by the time children reach the end of third grade. Our results point to an unintended consequence of a child care subsidy regime that conditions eligibility on parental employment and deemphasizes child care quality."
- Relevance to this filing specifically: the filing's recommended sequence is "fund retention, build supply, then extend demand-side eligibility." Extending demand-side eligibility is expanding exactly the work-conditioned, quality-deemphasising subsidy regime this paper studies. The filing scores that step on
passthrough_riskanddistributional. It has no dimension on which this finding could register at all.
2.4 The authoritative synthesis says the same thing
Duncan, Kalil, Mogstad & Rege, Investing in Early Childhood Development in Preschool and at Home, NBER WP 29985 (rev. Sept 2022; Handbook of the Economics of Education chapter). [Fetched: NBER]
This is not an advocacy source; Mogstad is co-author of the flagship pro-universal Norway papers. Their review splits exactly where this steelman splits:
- Ages 3–6: "For universal programs, the evidence is more consistent. For older preschoolers from disadvantaged families, this literature documents positive effects of preschool participation that persist into adulthood."
- Ages 0–3: "While evidence for older preschool-aged children suggests that participation in universal preschool is beneficial for disadvantaged children, the evidence on universal childcare for younger children is more mixed. In fact, some studies suggest that it has only limited effects on child development, or that it may even be detrimental."
- And they name the reason: "the counterfactual mode of care often differs between these two phases of early childhood," citing Bowlby's attachment phase (6 months to 2 years) as the theoretical prior.
- Their own summary table records Quebec as "Negative effects on children's health and noncognitive test scores. Negative long-term effects on adult health and life satisfaction, and higher crime rates," and Bologna as "Large and significant IQ loss for children of more affluent households, while small and insignificant for children of less affluent households."
Leg 3 — Where the gains are, and why "universal" is the wrong shape at 0–2
This is the leg that makes the case honest rather than adversarial: the benefits are real, and they are located.
3.1 The flagship pro-universal result, read distributionally
Havnes & Mogstad, Is Universal Child Care Leveling the Playing Field?, IZA DP 4978 → JPubE 127:100–114 (2015). [Fetched: docs.iza.org/dp4978.pdf]
The mean result from the same Norwegian reform (§3.4 below) is the single most-cited evidence that universal childcare works. Run on the earnings distribution with non-linear DiD:
- "While child care had a small and insignificant mean impact, effects were positive over the bulk of the earnings distribution, and sizable below the median."
- Verbatim from the results: "The estimated effects are zero or positive for all percentiles until the 69th, before turning negative at the upper part of the distribution … a decrease between 1 and 3.5 percentage points above the 77th percentile … The distribution is lifted by 4.6 percentage points at the median, by 5.9 percentage points at the 20th percentile … significant positive effects at every percentile between the 10th and the 60th, and significant negative effects from the 81st percentile."
- Mean impact: NOK −3,194, insignificant. Their conclusion: "mean impacts miss a lot."
- Their stated interpretation is my mechanism: "Children from disadvantaged families benefit the most from the child care reform, whereas it is less important or even detrimental for children from families with monetary and human capital to facilitate alternative arenas for child development of relatively high quality."
3.2 The children with the largest gains are the ones least likely to enrol
Cornelissen, Dustmann, Raute & Schönberg, Who Benefits from Universal Child Care? Estimating Marginal Returns to Early Child Care Attendance, IZA DP 11688 → JPE 126(6):2356–2409 (2018). [Fetched: docs.iza.org/dp11688.pdf]
- Identification: marginal treatment effects using a staggered municipal expansion in Lower Saxony, on the full population of compulsory school-entry examinations.
- Finding, verbatim: "children with lower (observed and unobserved) gains are more likely to select into child care than children with higher gains. This pattern of reverse selection on gains is driven by unobserved family background characteristics: children from disadvantaged backgrounds are less likely to attend child care than children from advantaged backgrounds but have larger treatment effects because of their worse outcome when not enrolled in child care."
- The policy inference is a targeting inference, and the paper makes it. If returns fall as you move up the enrolment margin, the marginal child added by universalising is close to the lowest-return child in the population. The design that captures the gains is one that reaches the non-attending disadvantaged — outreach, priority, intensity — not one that expands the offer to everyone.
- Scope note (against me): the German reform this paper studies mainly shifted enrolment from age 4 to age 3, not 0–2. I flag this rather than let it pass.
3.3 The best-identified 0–2 study in existence — and it splits the same way
Drange & Havnes, Child Care Before Age Two and the Development of Language and Numeracy: Evidence from a Lottery, IZA DP 8904 (Mar 2015) → JOLE 37(2) (2019), "Early Childcare and Cognitive Development." [Fetched: docs.iza.org/dp8904.pdf]
- Identification: an actual randomised assignment lottery for oversubscribed Oslo childcare places. Offer lowers starting age by ~4 months from a control mean of ~19 months; F on the instrument ≈ 100; covariates balanced.
- Mean effect is positive: lottery offer improves average performance at age 7 by ~12% of a standard deviation (language ~12% SD, mathematics ~same). IV: starting one month later costs just under 3% SD.
- But the heterogeneity is the finding. Verbatim: "Most strikingly, we find stronger effects among children from low income families. Indeed, among high income families, we find no impact of early child care start on performance in neither language nor mathematics. This suggests that child care policies may be more effective if targeted at low income households." The subgroup table shows family-income-high coefficients of −0.004 (language) and −0.004 (maths), both insignificant.
- Note the direction of the concession: this paper is counter-evidence to a blanket anti-0–2 claim and supporting evidence for the targeting claim. That is precisely why the steelman is stated as "universality is the wrong shape," not "0–2 care is harmful."
3.4 The mechanism, one more time, from the same authors
Havnes & Mogstad, No Child Left Behind: Universal Child Care and Children's Long-Run Outcomes, IZA DP 4561 → AEJ: Policy 3(2):97–129 (2011). [Fetched: docs.iza.org/dp4561.pdf] Strong positive effects on educational attainment and labour-market participation, reduced welfare dependency; "children with low educated mothers and girls benefit the most." Ages covered: 3–6 ("large variation in child care coverage for children 3–6 years").
And its companion, Money for Nothing? Universal Child Care and Maternal Employment, IZA DP 4504 → JPubE 95(11–12):1455–1465 (2011) [fetched]: "our precise and robust difference-in-differences estimates reveal that there is little, if any, causal effect of child care on maternal employment, despite a strong correlation. Instead of increasing mothers' labor supply, the new subsidized child care mostly crowds out informal child care arrangements, suggesting a significant net cost of the child care subsidies."
Leg 3, stated in one line: four independent, well-identified studies — Norway (distributional), Germany (MTE), Norway again (0–2 lottery), Italy (RDD) — converge on the same gradient. Returns are largest for disadvantaged children with weak home counterfactuals and fall to zero or below for advantaged children. Universality is the design that spends the most money on the lowest-return children, and at 0–2, on children for whom the return may be negative.
Leg 4 — Revealed preference: the filing reports the fact and reads it backwards
The filing's ws08-demand.md reports, from NSECE, that under-3s' modal arrangement is individual unpaid care (26.5%) and that ~40% of households with under-5s use no regular nonparental care. It draws one inference: "Universal provision is a behavioral offer to this group, not a subsidy to existing behavior — take-up assumptions drive everything." That is a cost-modelling inference. It never asks the prior question: is that 40% a group with an unmet need, or a group that has chosen something?
Evidence that at least part of it is choice:
- Over half of households with a child under five do not search for care at all. NSECE 2012 household brief, ACF/OPRE (Oct 2014) [fetched:
acf.gov/sites/default/files/documents/opre/brief_hh_search_and_perceptions_to_opre_10022014.pdf]: "Almost half (47 percent) of households reporting about a child under age five searched for care in the past 24 months." The complement — 53% — conducted no search in two years. Non-search is not proof of satisfaction, but it is the behaviour of a population that is not queuing. - Parents who do search for infant/toddler care are buying employment support, not child development. Same brief, verbatim: "Half (51 percent) of respondents searching for infant or toddler care did so for work reasons compared to just over one-quarter of those searching for preschooler care. Conversely, 41 percent of respondents reporting about preschoolers identified educational and social enrichment as the reason for a search compared to 19 percent of parents reporting about infants or toddlers." Four in five parents searching for infant care are not looking for a developmental intervention. The developmental case for universal 0–2 provision is one that parents of 0–2s do not themselves make.
- Stated preference points the same way. Pew Research Center, "Mothers and work: What's 'ideal'?" (19 Aug 2013, reporting a 2012 survey; question asked 1997/2007/2012) [fetched raw HTML]: 47% of mothers with a child under 18 say part-time work is their ideal, 32% full-time (44%/23% in 1997; 50%/20% in 2007). Married mothers: 53% part-time vs. 23% full-time. And on the on-point question — the ideal for women with young children — the public says 47% part time, 33% not working outside the home, 12% full time. A full-day, full-year, centre-based universal 0–2 entitlement is priced at a work pattern that a minority of mothers of young children name as their ideal.
- The gradient is where you would expect if this were partly preference and partly constraint: 40% of mothers with family income under $50,000 say full-time is ideal vs. 25% at $50,000+; 49% of unmarried mothers vs. 23% of married mothers. Which is also the population where the causal literature (Leg 3) says the child-development return is positive. Preference and return point at the same target group.
What would distinguish preference from unmet demand — stated plainly, because I cannot settle it here.
- A randomised offer with measured refusal. The Oslo design (Drange & Havnes) is the template: randomise offers of free or near-free places and record decline rates by family type. Take-up at zero price is the only clean measure of latent demand. Waitlists measure demand conditional on applying.
- A fungible-cash arm. Offer families the per-slot public cost as unconditional cash and observe how many still buy centre care. If care demand collapses when the money is fungible, the underlying demand was for income, not for care. Finland's and Norway's home-care allowances are the closest existing natural experiments — and
protocol-review.md§A2 explicitly told this filing to add a caregiver-choice architecture with the Nordic evidence attached. The architecture was added (a11); the Nordic evidence was not, andws14-architectures-v0.md:19still reads "Nordic maternal-employment cost still to be weighed." - Choice-set elicitation at zero price. NSECE asks parents to rate arrangement types on six characteristics; it does not ask which arrangement they would choose if all were free. That instrument does not exist nationally and would be cheap to add.
- Two-tier queue behaviour. Quebec's 30,688-child waitlist — the filing's one Quebec fact — is demand revelation, and it is evidence against me at the margin. The filing reads it purely as a supply failure. It is both.
And the honest limit on Leg 4, which I state here rather than in §3 because it is fatal to the strong version: the preference is highly price-elastic. From BGM's own long-run paper, citing Haeck et al.: between the mid-1990s and 2008 the share of Quebec children aged 1–4 in centre-based care as their primary arrangement rose "from under 10 percent to close to 60 percent," while the share in parental care fell "from around 55 percent to roughly 25 percent" — against a rest-of-Canada parental-care share that fell only from just under 60% to about 50%. Whatever American parents "prefer" at $13,000 a year is not what they will choose at $5 a day. Leg 4 therefore supports only the weak claim: the filing asserts unmet demand where it has measured non-use, and it never ran the test that would tell the difference. It does not support the strong claim that families would decline a free universal offer.
Leg 5 — The structural charge
Legs 1–4 are contestable empirics. Leg 5 is a claim about the filing's architecture and it is not contestable from inside the record:
- The filing's abstract commits to universal coverage 0–12 in its first sentence and asks only what it would take.
- Its pre-registered adjudication criteria contain no hypothesis whose "supported" branch means "do not do this."
- Its scorecard contains no dimension on which a negative child-development effect could be recorded, and its
child_dev_firstobjective operationalises child development as coverage of 0–2s and 3–4s. - Its Phase 0 successfully retrieved the three papers that would have raised the question and none reached any workstream.
- Its own protocol review predicted all of this in advance (§A3 objective function, §A4 "the hypothesis set tilts uniformly skeptical" — skeptical about means, not ends; §A5 missing economics literature; recommendation 8 adversarial presentation for Quebec child outcomes).
A filing may legitimately decide that whether to do a thing is out of scope and only how to do it is in scope. This one does not say so. It says "feasibility is a design and sequencing problem" in a whitepaper whose §7 reports negative RCT effects from a scaled pre-K programme, and whose honesty box lists nine acknowledged weaknesses, none of which is "we assumed the goal."
3. The strongest counter-evidence against my own case, stated fairly
If I were defending the filing, this is what I would use, and some of it is very strong.
C1 — The most recent and most direct rebuttal: Quebec's negative childhood effects do not show up in adult economic outcomes, and the programme's welfare return is high. Montpetit, Carrer & Beauregard, A Welfare Analysis of Universal Childcare: Lessons From a Canadian Reform, working paper dated 9 April 2026 [fetched: sebastienmontpetit.github.io/WebsiteSM/MCB_QCchildcare.pdf]. Structural model of childcare demand + MVPF framework on the same 1997 reform, using the same BGM NLSCY sample definition, plus Canadian RDC data.
- "We find that the negative effects on behavioral outcomes in childhood documented by Baker et al. (2008, 2019) do not translate into lower educational attainment or reduced earnings in their early career." Event-study estimates on eligible cohorts' labour income are insignificant positive. They state: "we can rule out large negative effects on children's long-run economic outcomes."
- They price BGM's crime finding: "Given the nature of typical juvenile crimes, these additional social costs turn out to be relatively small compared to mothers' gains in this context."
- MVPF 2.83 (90% CI [2.02, …]); benchmark MVPF from mothers' earnings alone 1.26. Mothers of preschoolers in Quebec earned +$3,003/yr (constant 1997 dollars) relative to control.
- They also find the supply channel matters more than the price channel — which is the filing's central thesis, arrived at independently.
- This is the single strongest strike against me. It says: the harm is real in childhood, it does not compound, and the programme was worth doing anyway.
C2 — BGM's crime finding was publicly contested by criminologists. Jane Sprott (Ryerson) called the child-care/crime link "absurd"; Ronald-Frans Melchers (Ottawa) said "substantial changes in the processing of young offenders over these years affected Quebec and the rest of the country very differently" [Globe and Mail, 2015; fetched via page summariser — press tier, and I did not read the criminologists' own writing]. In fairness to BGM: the objection was aimed at the 2015 working paper, and the published AEJ:Policy version restricts crime data to 2006+ expressly to clear the YCJA transition and reports both accusations and convictions. That answers the timing objection. It does not answer the deeper one — the estimand is an intent-to-treat over all eligible Quebec children, not an effect on attendees.
C3 — The most careful replication returns a mixed, not a confirming, verdict.
- Kottelenberg & Lehrer, NBER WP 18785: while confirming BGM's robustness, they add — verbatim — "despite estimated effects stemming from the policy indicating declines in motor-social development scores in Quebec relative to the rest of Canada, our analyses imply that on average attending childcare in Canada leads to a significant increase in this test score," and that "most of the negative impacts reported in earlier research are driven by children from families who only attended childcare in response to the implementation of this policy."
- Kottelenberg & Lehrer, Targeted or Universal Coverage?, NBER WP 22126 → JOLE 35(3):609–653 (2017) [fetched]: using Athey–Imbens change-in-changes, "the Quebec Family Policy significantly boosts developmental test scores for children from single parent households particularly for those who are most disadvantaged and located at the lower quantiles of the distribution. However, children from two-parent families between the 10th and 50th quantile generally receive significant negative impacts from child care … those in the top half of the distribution are generally unaffected by the policy."
- Note carefully what this does to my Leg 3: it supports the targeting story but inverts my gradient in one place — here the losses land in the lower-middle of the two-parent distribution and the top is unaffected, whereas Havnes–Mogstad and Fort–Ichino–Zanella put the losses at the top. The gradient is real; its exact location is not settled.
C4 — The parenting damage is temporary. Haeck, Lebihan, Lefebvre & Merrigan (2015): "However, negative effects of the program on parental behaviors vanish when the child is in school. Moreover, we find that this pattern persists even ten years after the implementation of the reform."
C5 — Montpetit et al. also flag two dosage-based challenges to the long-run findings: "evidence from Haeck et al. (2018) and Ding et al. (2021) challenge some of these findings, showing that accounting for treatment dosage reduces the magnitude of long-term effects." — UNVERIFIED beyond this citing sentence; I did not fetch either paper.
C6 — The one US universal precedent is strongly positive. Herbst, Universal Child Care, Maternal Employment, and Children's Long-Run Outcomes: Evidence from the U.S. Lanham Act of 1940, IZA DP 7846 → JOLE 35(2):519–564 (2017) [fetched]: "the Lanham Act had strong and persistent positive effects on well-being, equivalent to a 0.36 standard deviation increase in a summary index of adult outcomes … the benefits of the Lanham Act accrued largely to the most economically disadvantaged adults." A heavily-subsidised, universal, US programme covering children including under-5s. Universality per se is not the culprit. (The filing lists Lanham in its anchor table as "not yet verified — queued," so it does not use this either.)
C7 — Randomised US evidence at 0–3 is positive. Early Head Start Research and Evaluation Project (ACF/OPRE; Mathematica + Columbia CCF): 3,001 families, 17 sites, random assignment, assessments at 14/24/36 months [fetched: acf.gov/sites/default/files/documents/opre/research_brief_overall.pdf]. At age 3, programme children scored 91.4 vs. 89.9 on the Bayley MDI and 83.3 vs. 81.1 on the PPVT, were significantly less likely to score in the at-risk range, engaged parents more and were rated less aggressive; parents were more emotionally supportive, read daily more often (56.8% vs. 52.0%) and spanked less (46.7% vs. 53.8%). The brief is candid: "these overall impacts were generally modest in size." This is the best US causal evidence for 0–3 non-parental care — and it is targeted, comprehensive, and two-generation, which is the shape my steelman argues for, not the shape the filing proposes.
C8 — The counterfactual argument cuts against blanket opposition in the US context. Fort–Ichino–Zanella's own model says the sign flips with home-care quality. In the United States, a large share of the 0–2 counterfactual is unregulated informal care — which is why the filing's own §8 finding (26.5% of under-3s in individual unpaid care) is ambiguous evidence: the same mechanism that predicts harm for affluent Bolognese children predicts gains for many American children. My case therefore cannot be "universal 0–2 care harms children"; it can only be "the return varies by counterfactual, and universality is indifferent to the counterfactual."
C9 — My Leg 4 is the weakest leg and I said so above. Preference at $13k is not preference at $0. Quebec 10%→60%.
C10 — Every long-run Quebec study is a DiD against "the rest of Canada." The programme also coincided with other Quebec family-policy changes (BGM discuss these; Montpetit et al. note the 2006 Quebec Parental Insurance Plan is a separate large shock and drop NLSCY cycles for that reason). None of these designs is a lottery. Drange & Havnes is — and it is the one that comes out positive.
4. Adjudication (S4) and scorecard sensitivity (S5)
4.1 Verdict: the steelman partially survives.
It does not win on the policy. C1 is dispositive against the strong claim. The best current estimate of Quebec's overall welfare effect is strongly positive (MVPF 2.83), and the childhood harms do not appear in adult education or earnings. C6 and C7 show universal and 0–3 public care can produce large, durable, pro-poor gains. Anyone claiming from this literature that universal 0–2 care is net harmful is overreading it.
It wins decisively on the framing. Three findings stand and are not answerable from inside the filing's record:
- The filing's central question was never asked, and its instruments cannot ask it. No hypothesis, no adjudication criterion, and no scorecard dimension can return "the goal is wrong." The
child_dev_firstobjective defines child development as 0–2 and 3–4 coverage. This is not a gap in evidence; it is a property of the apparatus. - The relevant literature was retrieved and dropped.
phase0-findings.md:38names BGM 2008, BGM 2019 and Cornelissen et al. 2018 as successfully retrieved. None reached a workstream.protocol-review.md§A5 predicted the omission and recommendation 8 required adversarial presentation of Quebec child outcomes specifically. The filing ran that discipline for Tennessee vs. Boston and not for Quebec. Under the filing's own §0.1 adversarial-presentation rule, this is a protocol deviation that never made the deviations log. - The sign of the 0–2 effect is conditional on the counterfactual and on family background, across four independent identification strategies (non-linear DiD in Norway; MTE on the full German school-entry population; a randomised lottery in Oslo; RDD on admission thresholds in Bologna) — and the field's authoritative review (Duncan/Kalil/Mogstad/Rege) states plainly that for 0–3 the evidence is "more mixed … may even be detrimental." A filing whose recommendation reaches 11.1 million under-3s cannot leave this out of the record.
Required consequences under S4 ("apply it to the record, not to a footnote"):
- Declare the objective. The filing must state, in the whitepaper, that it assesses feasibility of a stipulated goal and does not adjudicate desirability — or else adjudicate it. Right now it does neither and reads as if it had.
- Add the age-band scope qualifier. "Universal coverage 0–12" should not be stated as an undifferentiated goal. The evidence supports universality most strongly at 3–4 and 5–12 and least at 0–2, where the strongest finding is heterogeneous returns, not a mean benefit.
- Open a
ws07Quebec section with the adversarial-pair treatment the filing's own protocol requires: BGM 2008/2019 and Fort–Ichino–Zanella against Montpetit et al. 2026, Kottelenberg–Lehrer 2013/2017 and Drange–Havnes. - Log the deviation. Recommendation 8 of
protocol-review.mdwas applied selectively. That belongs indeviations-log.md. - Add a Known-Unknowns entry for the preference-vs-constraint test at 0–2, with the four discriminating designs listed in Leg 4.
The most interesting result of the whole pass: the steelman largely vindicates the filing's 0–2 architecture while destroying its justification. ws14/ws16 recommend, for 0–2, a supply-first plus caregiver-choice blend rather than a centre-led universal build — reached from revealed preference and supply economics. The child-outcome literature independently supports the same destination, by a completely different route (returns are highest where the home counterfactual is weakest; paid kin and home-based care keep the adult-to-child ratio near 1). The filing got the right answer for 0–2 without the argument that most strongly supports it. That is worth saying in the record, because a right answer resting on the wrong reason is fragile to exactly the objection the filing never heard.
4.2 S5 — scorecard sensitivity
The contested cell is not a cell. It is coverage_02 (scale: 1 = "reaches almost no 0–2s"; 5 = "universal 0–2 offer including home-based settings") and its weight of 2 inside child_dev_first. Re-running rank.py with the steelman's reading substituted:
| Reading | Top 4 |
|---|---|
child_dev_first as filed |
a3 3.90 · a8 3.85 · a6 3.75 · a9 3.60 |
A. coverage_02 weight 0 (agnostic on whether 0–2 coverage helps children) |
a3 3.89 · a8 3.83 · a6 3.83 · a9 3.56 |
B. coverage_02 inverted, weight 2 |
a6 3.75 · a3 3.70 · a8 3.65 · a4 3.60 |
C. inverted + preference_fit weight 3 |
a6 3.71 · a3 3.67 · a8 3.57 · a4 3.52 |
- Under A (the defensible minimum — merely declining to assume 0–2 coverage is a child-development good) the top-4 set holds and the ordering barely moves. The filing's recommendation survives agnosticism.
- Under B/C the top-4 set changes: a9 (federal fallback ladder) drops out and a4 (K–12 extension) enters, and a6 (supply-first) takes first place under the child-development objective for the first time.
- But I do not offer B/C as the right re-score, and here is why — it is the sharper S5 finding. Inverting
coverage_02also demotes a11 (caregiver-choice allowance) from 5 to 1, and a11 is the architecture the steelman's own evidence most supports: paid kin and home-based care preserve a near-1:1 adult-to-child ratio, which is exactly the counterfactual Fort–Ichino–Zanella identify as the source of the advantage. The inversion mangles a11 becausecoverage_02does not distinguish centre-based coverage from parental/kin coverage. The scorecard has no dimension capable of expressing the contested proposition, and any attempt to encode it corrupts a dimension that means something else.
S5 conclusion: the fix is not a re-weighting. It is a new dimension — something like child_effect_0_2 anchored on the counterfactual-quality mechanism (1 = shifts children with strong home counterfactuals into full-time centre care; 5 = concentrates intensity on children with weak home counterfactuals, or funds near-1:1 settings) — plus a re-score of all twelve architectures against it. Until that exists, the filing's child_dev_first ranking should be reported as what it is: a ranking on coverage and workforce inputs, relabelled.
5. Honest self-assessment
What I am confident in.
- The framing finding (§1, Leg 5, §4.1 items 1–2). It is verifiable from the repository in minutes and does not depend on adjudicating any empirical dispute.
rank.py'schild_dev_firstweights andscales.md's 14 dimensions are the strongest single piece of evidence in this document, and they are the filing's own files. - The existence and seriousness of the Quebec outcome literature, and that the filing retrieved and dropped it.
- The heterogeneity result (Leg 3). Four independent identification strategies, three of them quasi-experimental at population scale and one an actual lottery, converging on "returns fall as the home counterfactual improves." Duncan/Kalil/Mogstad/Rege state it as the field's position.
What I am not confident in.
- That universal 0–2 care is net harmful. I do not believe the evidence shows this and I have not argued it. C1 (Montpetit et al. 2026, MVPF 2.83, no adult economic penalty) is the best current estimate and it goes the other way. If the reader takes away "universal childcare hurts kids," I have failed at the assignment.
- Leg 4 (revealed preference) is the weak leg. Non-search is not refusal; Quebec's 10%→60% shift shows the preference is price-elastic; the Pew data are 2012 and are stated, not revealed. The most I can honestly claim is that the filing asserted unmet demand where it had measured non-use, and never specified the test. I listed four tests that would settle it and can run none of them here.
- The exact location of the negative region is unsettled. Havnes–Mogstad put it above the 69th–81st percentile; Fort–Ichino–Zanella have it rising with income; Kottelenberg–Lehrer put it in the 10th–50th percentile of two-parent families with the top unaffected. These cannot all be right. I have reported the disagreement rather than picking the one that suits me.
Method limitations to record.
- OpenAlex was rate-limited (paid tier, $0 balance) for this entire session, so I could not run systematic citation sweeps. My literature identification is therefore seeded from named papers plus reference-chasing through Montpetit et al. and Duncan et al. It is not a PRISMA-style search and I have not established that I found the best contrary evidence — only good contrary evidence.
- Three items rest on materials I did not read in raw form and are marked accordingly: the Globe and Mail criminologist quotes (press tier, via page summariser); the Haeck et al. 2018 / Ding et al. 2021 dosage challenges (UNVERIFIED — known only from Montpetit et al.'s citing sentence); Drange & Havnes' published JOLE effect sizes (0.16 SD language / 0.11 SD maths / 0.26 SD low-income) which I read only inside Duncan et al.'s summary table — the numbers I quote in Leg 3.3 (12% SD, ~3% SD per month) are from the IZA working paper I read directly.
- Fort–Ichino–Zanella was read from a University of Chicago Press manuscript preprint (DOI-stamped) hosted on a third-party mirror, not from the JPE site. Content and DOI match; the host does not.
- I did not verify the filing's own NSECE figures (26.5%, ~40%); Phase 1 covered them and I treat them as given.
pdftotextwas unavailable in this environment; all PDF extraction used PyMuPDF.- I did glimpse
childcare/research/steelman-log.mdwhile grepping the record — an orchestration file naming this target. It confirmed the assignment and contributed nothing to the argument; everything in §1 was derived from the filing's own published files.
On the S3 test — "would this persuade someone who currently holds the filing's conclusion?" Partly, and in a specific direction. It would not persuade them to oppose universal childcare. It should persuade them that (a) their filing answered a question it never asked itself whether to ask, (b) the age band where 64% of the difficulty and 100% of the developmental controversy live was assessed on coverage counts alone, and (c) their own 0–2 recommendation is better supported than they know, by a literature they retrieved and discarded. That is a qualification and a scope narrowing, not a withdrawal — S4's middle outcome.