GUBMENTPlain talk · policy frontier
Filings / Elder care / Sources & data
Sources & Data · Series GBMT-7 · Filed 2026-08-03

The research record behind Whitepaper No. 7

Every verified anchor, every workstream's finding, and every logged deviation from the protocol — the full record behind "Elder care: does workforce or funding actually cap home care?"

Record 0

Protocol & scope

Long-term services and supports (LTSS) for older Americans and adults with disabilities — home- and community-based care, assisted living/residential care, skilled nursing, unpaid family caregiving, and the financing that sits under all four (Medicaid, Medicare's narrow post-acute role, private long-term care insurance, and out-of-pocket spend-down). Imports the gubment method (M1–M8) in full: two-source rule, root-tracing for every claim, pre-registered hypotheses, and a deviations log that stays visible rather than getting edited away.

Phase 0 verdict: GO

Phase 0 (2026-08-03) verified the anchor table's starred rows and committed a first ACS-derived population baseline (65+ and 85+ by state) as an auditable data pipeline. Its own kill condition fired as a headline: the unpaid family-caregiving valuation cannot be reconciled to a defensible single range across independent sources. Two more anchors broke or reframed outright — the 2024 CMS nursing-home minimum-staffing rule turned out to be dead, not merely contested, and assisted living does not out-census skilled nursing after all. Eleven workstreams (§2–§12) were then run to varying depth, followed by a full red-team pass (Workflow, 11 agents: four attack/defend pairs, a re-score, reconciliation, and whitepaper synthesis) that resolved 22 items and reconciled 3 scorecard cells before publication. That re-score was instructed not to read the scorecard rather than structurally prevented from reaching it; a genuinely blinded pass ran on 2026-08-10 and moved 25 of the scorecard's 40 ranked cells, withdrawing one headline finding. See Record 4 and ws11-rescore-log.md. Two independent verification passes then ran against this filing on 2026-08-10 under the Verification Protocol: a Phase 1 fact-check against primary statute, rule and paper text (689 claims, 40 corrected, 42 reworded), and a Phase 2 steelman against this filing's verdict, its top-ranked architecture and its scorecard headline (Record 5), which withdrew two more published claims.

Record 1

Anchor table — priors, stated before evidence, verified after

Every anchor below was written down as an unverified guess before research began, so it could be broken. ★ rows were verified in Phase 0; the rest were verified during full execution or remain queued.

#Anchor (unverified prior)Verified value & delta
1 ★A majority turning 65 will need some paid or unpaid LTSS before deathASPE 2022 (current, DYNASIM4): 56% will develop severe need. ASPE 2019 (HRS-based): 70%. Kemper/Komisar/Alecxih 2005 (independent model): 69%. Majority claim holds either way. Softened on Phase 1 verification (2026-08-10): both figures are ASPE's, and the 2022 brief does not supersede the 2019 one — different model (DYNASIM4 vs. HRS) and different construct, with the 2019 study still cited as live. Cite 56% as the current primary figure and name the model; "superseded vintage" and "citation trap" overstated it.
2 ★Direct-care workforce needs 1M+ additional jobs by early 2030sPHI 2025 (2024–2034, BLS-derived): 772K+ new jobs from growth; 9.7M total openings including replacement/turnover. Corrected on red-team review: the original "~10x undersold" comparison mixed growth-only against growth+replacement. Like-for-like (772K vs. "1M+") is ~25% undersold, not 10x — the 10x framing is retired.
3 ★Insurers actively selling new individual private LTC policies fell from 100+ to a small handfulPeak 125 (2000)→104 (2002), AHIP survey, never repeated. ~12–15 by 2014 (NAIC). Current ~7–10 standalone individual carriers (2024–25 trade lists); LIMRA corroborates >75% exit by 2012. Direction/magnitude confirmed, but no rigorous market census exists post-2016; the market has partly shifted to hybrid life/LTC combo products ($4.2B new premium, 2024).
4 (kill condition test)Unpaid caregivers number in the tens of millions, imputed value in the hundreds of billions annuallyAARP/NAC 2025: 63M caregivers (all ages). AARP Valuing the Invaluable 2026 (2024 data): 59M adult caregivers, 49.5B hours, $1.01T — up from $600B (2021) and $350B (2006). CBO 2013 (elderly-only, replacement-cost): $234B. RAND 2014 (elderly-only, opportunity-cost): $522B. Not reconcilable to a single defensible figure — the headline finding. Red team clarified this as three named methodological drivers (population scope, wage-rate basis, hours source) explaining the spread, not true irreconcilability — but the instruction to report a conditional range, never a point estimate, stands.
5Most states impose a Medicaid countable-asset limit near $2,000 for an individual applicantNot attempted in Phase 0 and not covered in this pass's workstream findings — queued for future execution.
6 ★WA Cares payroll-tax rate and lifetime benefit cap as enactedRate 0.58% of gross wages, no wage cap. Corrected on Phase 1 verification (2026-08-10): RCW 50B.04.080(1) makes .58% the initial rate and provides that from January 1, 2026 the pension funding council sets it biennially, capped at .58% — so “unchanged since enactment” no longer describes the statute. Cap $36,500 (2026), auto-inflation-adjusted. Premiums delayed Jan 2022→July 2023; opt-out window through Dec 2022; SB 5291 (2025) loosened vesting, no rate/cap change. Benefits first became payable July 1, 2026. "Delayed and revised once" undersold the amendment history — most consequential, this research lands weeks after the program's first-ever payout, with essentially zero months of real payout data yet.
7 ★The 2024 CMS nursing-home minimum-staffing rule sets an HPRD floor with a phased timelineOriginal (Apr 2024): 3.48 total HPRD, 0.55 RN, 2.45 NA, 24/7 on-site RN, phased 2026–2029. Rescinded effective Feb 2, 2026 (CMS interim final rule, Dec 3, 2025) following vacatur in two federal courts plus a statutory moratorium via H.R. 1 (2025). Corrected on Phase 1 verification (2026-08-10): Pub. L. 119-21 §71111 runs from enactment (2025-07-04) to September 30, 2034 — 9 years 3 months, not ten — and CMS’s own repeal rule says “September 30, 2034” throughout, not the 2035 this row previously attributed to it. Facility-assessment and Medicaid staffing-spend-transparency provisions survive. MAJOR reframe: not "litigation and rescission risk," the rule is dead by both courts and Congress — changes the research question from assessing an emerging standard to assessing the aftermath of a reversed one.
8Medicaid HCBS waiver waiting lists sum to roughly 500,000+ nationally, with substantial double-countingKFF tracker: ~692,679 nationally in 2021, growing to ~710,000+ by 2024 (+2.6% 2023→2024); average wait fell 45→36 months (2021–23) then rose back to ~40 (2024). CA, NM, TX together hold over half of all waitlisted individuals. Texas's own 2015 data shows a ~95% duplicate-listing inflation. Seed undersold the scale — actual figure ~40% higher than the seeded 500,000+ floor, still rising through 2024 despite ARPA §9817's targeted spending.
9Nursing homes are a minority of paid LTSS spending; HCBS has grown to be the larger shareNot attempted in Phase 0; confirmed during full execution — of KFF's $415B (2022) compiled LTSS total, HCBS accounts for $284B vs. institutional $131B; HCBS spending crossed institutional spending nationally in 2013.
10PE-owned nursing homes show a measurable mortality increase vs. non-PE facilities in a widely cited studyGupta, Howell, Yannelis & Gupta, Review of Financial Studies 37(4) 2024 (NBER WP 28474): 4.2M unique short-stay Medicare patients, 2005–2017, within-facility IV design. OLS: +0.3 pp, ~2% of the mean; causal estimate: +11% (~22,500 excess deaths and ~172,400 lost life-years over the sample). Also +8% billed spending per stay (+6% including the following 90 days). Corrected at primary tier on Phase 1 verification (2026-08-10): this row previously read “>7M patients,” “OLS: +10%,” “~20,150–21,000 excess deaths,” “+19% billed spending” and “~50% higher antipsychotic use.” The paper’s full text supports none of the five — there is no antipsychotic result in it at any vintage, and 20,150 is the superseded 2021 draft’s figure. The +11% is a local average treatment effect for the patients steered into PE homes by distance; the authors report “small beneficial effects for some patients” and conclude the harm falls on a subset. Headline estimate confirmed — one JAMA COVID-era comparison found PE facilities in line with industry average, which the study's own co-author excluded as inconclusive; confidence downgraded from high to medium-high on red-team review, core estimate unrebutted.
11 ★Assisted living serves a comparable or larger number of paid LTSS recipients than skilled nursingAL/residential care: ~1.016M residents (CDC/NCHS NPALS, 2022). SNF: ~1.24M residents (KFF/CMS CASPER, July 2025). Corrected on red-team review: prior likely broken on current data, but vintage-mismatched (2022 vs. 2025) and directionally unstable — do not treat the ~20% SNF-larger margin as current; no independent current AL count was obtainable (NIC MAP paywalled). The research/regulatory-attention-asymmetry half of the claim remains plausible.
12OAA Title III-E (NFCSP) appropriation is roughly two orders of magnitude below Medicaid HCBS spendingNFCSP FY2026 appropriation ≈ $209M (ACL). Medicaid HCBS spending ≈ $116B (FY2020) to $162B (CY2020); total Medicaid LTSS ≈ $257B (2023). Magnitude confirmed but undersold — NFCSP is ~three orders of magnitude below Medicaid HCBS ($209M vs. ~$140B), not two. Against row 4's caregiver-value range, NFCSP funds well under 0.1% of even the lowest independent valuation.
Record 2

Workstream findings

Eleven workstreams outlined in the protocol; nine produced dedicated findings files, each executed with a full source register and two-source-rule discipline. Full writeups (source-by-source, with confidence ratings) are in each workstream's own findings file in the repository.

§2 · Baseline & spendingNo single canonical "LTSS spending" number exists — every figure is a downstream analyst choice

Best available compiled total is KFF's $415B (2022): Medicaid 61% (~$253B), out-of-pocket 17% (~$71B), Medicare + private LTC insurance combined ~21% (~$87B, not separable by any current source found). By setting, HCBS ($284B) exceeds institutional spending ($131B). Scope corrected on Phase 1 verification (2026-08-10): the 2013 crossover this line previously attached to the all-payer total is a Medicaid LTSS event; KFF dates no all-payer crossover. Independent models (Urban Institute DYNASIM, CBO) diverge substantially from NHEA-based accounting — the field has no single measurement convention, and every "LTSS spending" figure embeds an unstated choice about which claims lines count.

§3 · WorkforceMoney went to wages, wages rose, and neither vacancies nor waiting lists improved — but a wage-inflation confound leaves it INDETERMINATE

ARPA §9817 put 71% ($26.3B of $37.1B) toward workforce recruitment/retention; wages moved substantially where states spent it that way (Colorado's direct-care wage $12.41→~$18), yet every state responding to KFF's 2023 survey still reported shortages and national HCBS waiting lists grew from ~692,679 (2021) to ~710,000+ (2024). Red team downgraded the verdict to plain INDETERMINATE, since economy-wide post-pandemic wage inflation is an unruled-out confound. Separately, the direct-care workforce's foreign-born share (28% nationally, 36.5% for home health aides on 2019 census data — roughly twice the general workforce; “more than double” held against the 2019 base of 17% but not against 2024’s 19.2%) is a well-documented exposure to immigration-enforcement risk, though the actual labor-supply effect is anecdotal-only so far.

§4 · Financing & private LTC insuranceOne state modelled reinsurance and the study found little savings — and private LTC insurance now sells only to the wealthy

Washington's 2015 Milliman feasibility study found a public-private reinsurance option had "little potential to generate savings" and couldn't revive the private market — the same study that led Washington to build WA Cares instead. Reinsurance shifts catastrophic-tail risk without touching the adverse-selection mechanism that killed the CLASS Act. Hybrid life/LTC combo products, the dominant new-sales channel since ~2014, are pitched at buyers with $500K–$5M net worth — sidestepping adverse selection by selling into an already-wealthy, self-selected pool rather than fixing it.

§5 · HCBS waiting listsIndiana is consistent with the "architecture, not money" hypothesis (n=1, correlational, confounded); Texas, the state most cited for it, can't be tested at all

Indiana's combined aged/disabled waiver waitlist grew from ~12,800 (2024) to over 17,000 (Feb 2026) atop a spending mix sending roughly two-thirds to three-quarters of Medicaid LTSS dollars to institutional care (34% HCBS per the state's own report, ~23% per AARP's Scorecard) — well below the ~53% national average. Texas, the state most commonly cited for the institutional-bias thesis, turned out to be the one where the thesis is least verifiable: excluded from the two leading national trackers for data-accuracy reasons, with its own 2015 data showing a ~95% duplicate-listing inflation in raw waitlist counts.

§6/§9 · Institutional quality & federalismThe federal staffing rule is dead twice over; state law is the more durable lever, but enforcement lags

The 2024 CMS minimum-staffing rule was vacated by two federal courts and hit with a 10-year congressional moratorium before CMS formally rescinded it, effective February 2026. 38 states + DC already exceeded the old federal floor, and a Health Affairs 22-state panel (March 3, 2026 issue; previously dated 2025 here) found staffing mandates raised direct-care staffing ~5% with no effect on facility closures — undercutting industry's core objection, for the range of mandates studied. But New York's mandate shows statutory survival isn't enforcement: ~400 violating facilities went years with essentially no penalties. Separately, private-equity ownership of nursing homes shows a precise, largely uncontested +11% causal mortality effect (Gupta, Howell, Yannelis & Gupta, Review of Financial Studies 2024, >7M Medicare patients).

§7 · Assisted livingTwo GAO reports, eight years apart, found two different undercounts — and the incident-reporting fix doesn't take effect until 2027

GAO-18-179 (2018) found 26 of 48 studied states couldn't report critical-incident counts in Medicaid-funded assisted living; GAO-26-107884, Assisted Living Facilities: Information on Federal Spending and Medicaid Coverage (June 2, 2026), documents a different federal blind spot — "at least $12 billion" of Medicare and Medicaid spending in the setting in 2024, "likely an undercount because of data limitations." Corrected on Phase 1 verification (2026-08-10): this digest previously read the 2026 report as confirming the incident-reporting gap. It makes no incident-reporting finding. Two reports, two undercounts, of different things. A CMS rule requiring annual state incident reporting doesn't take effect until 2027 at the earliest. Medicaid's services-only AL footprint (~18–20% of residents) sits inside the broader HCBS-over-institutional spending shift since 2013, but no clean time series exists to independently quantify AL's own growth — directionally supported, precisely unquantified.

§8 · Family caregivingThe valuation figure can't be pinned to one number, but the health and labor-market evidence is solid and getting worse

AARP/NAC's own headline caregiver-value figure moved $350B (2006)→$600B (2021)→$1.01T (2024), and two independent estimates for a comparable era (CBO $234B, RAND $522B) diverge from it and each other by more than 2× on scope, wage method, and hours source. The Cash & Counseling Demonstration (1998–2003) is one of the only true RCTs in the entire elder-care domain: treatment caregivers reported less strain and better health, and the mechanism was closing an unmet-need gap, not efficiency — though costs ran higher in year one (AR +17%, FL +14%). CDC's BRFSS data shows caregivers' mental-health gap versus non-caregivers widened, not narrowed, from 2015–16 to 2021–22.

§10 · International precedentsGermany, Japan, and the Netherlands are all under real current strain — not stable models to import wholesale

All three systems show live cost and workforce pressure inside the measurement window, not just projected future strain: Germany's contribution rate rose again for 2025 and needed federal loans to hold 2026 rates steady, (a “110% rise in vacancies against 45% demand growth” previously stated here is withdrawn on Phase 1 verification, 2026-08-10 — it was cited to NBER w31870, which contains no vacancy or demand series); Japan projects a 570,000-worker shortfall by 2040 and is already proposing copay increases; the Netherlands projects a 266,000-worker shortfall by 2035. Japan's favorable reputation traces substantially to one clustered research lineage — the same single-lineage risk pattern housing found with Auckland. Each is better scored as a design-fragment donor than a whole-system precedent to import.

§11 · Architecture scorecardCash & Counseling is the only architecture with real trial evidence — and a second blind re-score found most of the rest of the board was measuring our own research effort

Nine candidate architectures scored on five anchored objectives (cost containment, institutional quality, workforce sustainability, caregiver relief, aging-in-place), across three scoring passes. Cash & Counseling is the only architecture backed by an actual RCT and wins decisively under caregiver-relief weighting — but its own RCT documented a real cost increase, which corrected its cost-containment score from 3 to 1, so it no longer wins under every weighting as the first pass claimed. A third pass, structurally blinded, moved 25 of the 40 ranked cells and exposed a defect in the anchored scale itself: high scores required cited evidence while low scores accepted its absence, so under-researched instruments were pushed to the floor. Twenty cells at 1 or 2 rested on no evidence in either direction; corrected, 30 of 40 sit at neutral. The claim that private LTC insurance market reform "ranks last under all three weightings, with no exceptions — the clearest, most stable negative finding in the whole scorecard" is withdrawn: four of that row's five cells were empty, and its last place survives only under the unamended scale. The direct-care wage floor also loses its lead on workforce weighting under both readings, tying the unstructured-cash comparator. Phase 2 steelman (2026-08-10): the replacement headline — federal LTC social insurance "ranks last under all three weightings … on evidence" — is itself a one-cell finding. Neutralise the single workforce cell that carries it (marked down on Japanese and Dutch national workforce shortfalls, which the United States shares without having the instrument) and the row ties for last under all three weightings rather than holding it alone. Published as a sensitivity, not entered. The same pass tightened this row's HCBS basis: the paper cited for the board's strongest positive cell also reports a null for Community First Choice, which this page had not carried.

Record 3

Deviations log

Every departure from the protocol as originally written, logged with its reason and effect on findings — per the method, this stays part of the record, not an appendix to it. 37 entries — the count was corrected on Phase 1 verification (2026-08-10), which also found that entries 22–29 (the PACE late-pass sequence) had never been published here at all; entries 35–37 are the Phase 2 steelman. The digest below still shows a condensed selection; the complete table is in the repository and the full list is now in the deviations log page.

#DeviationEffect on findings
1Anchor verification for rows 1, 2, 3, 6, 7, 11 delegated to six parallel subagents, QA-reviewed by the primary session with two independent spot-checks on the two highest-stakes claimsNo figure entered unverified; rows 6 and 7 independently confirmed against separate outlets
2Row 7's moratorium end-year is unreconciled: CMS's own repeal cites 2035, one law-firm summary states 2034Both figures recorded rather than picking one; doesn't affect the row's headline finding (rescinded Feb 2, 2026)
3Anchor rows 4, 5, 8, 9, 10, 12 not attempted in Phase 0Left blank, not guessed; all but row 5 closed during full execution
4, 16§2 baseline pulled population (ACS) only in Phase 0; even after full execution's spending pull, Medicare-alone and private-LTC-insurance-alone shares of the $415B KFF total remain unsplit by any source foundPopulation leg committed and auditable; the spending leg's payer-split gap recorded as a named data gap, not guessed at — third instance this session of "the US doesn't measure this domain the way its debate assumes"
5, 6§12's precedent scan (WA Cares + CLASS Act + one HCBS waiting-list state case) delegated to subagents; the Texas case study could not verify a current STAR+PLUS headcount or the state's institutional-vs-HCBS spending mix (HHSC pages returned HTTP 403; Texas is excluded from two national trackers)Both reports disclosed their own confidence gaps; "cannot confirm whether Texas is institutionally biased" entered as the finding for that half of the case study
7§3's H3.1/H3.2 were tested without pre-registered M3 adjudication criteria (refuted-if/supported-if/indeterminate-if)Verdicts reported conservatively (INDETERMINATE for H3.1; split exposure/effect verdict for H3.2) rather than retrofitting criteria to match the result
8Anchor row 8 (HCBS waiting-list total) filled in from data surfaced incidentally during §3's ARPA §9817 research, not a dedicated §5 passFilled opportunistically rather than left blank; §5 still needed its own dedicated pass for the duplicate-counting question
9, 10§8's Credit for Caring Act has no CBO/JCT score (never reached committee markup); the self-direction-vs-agency-workforce interaction (linking §3 and §8) was not resolved — the one source that raises it explicitly calls it untestedBoth recorded as explicit non-findings rather than filled with unsourced guesses or inference dressed as evidence
11, 12§6/§9's state staffing-mandate HPRD list is sourced from a legal/secondary aggregator, not checked against each state's statutory text; the enforcement-durability finding rests on a single state case (New York)Recorded with the caveat stated rather than presented as admin-data-verified or general; whether NY's enforcement gap generalizes is queued
13§5's Indiana test found one of two spending-split figures (AARP Scorecard, ~23% HCBS) could not be re-verified against AARP's primary page directly; the state's own report figure (34%) was independently confirmedBoth figures recorded with the discrepancy stated; the directional finding (well below the 53% national average either way) doesn't depend on which is exact
14§11's scorecard began as a single-scorer first pass, with PACE and private-LTC-insurance market reform initially left unscoredRecorded explicitly rather than presented as final; a full red-team re-score (see Record 4) later reconciled 3 cells, though PACE and Credit for Caring remain unscored placeholders
22Late PACE/§12 evidence pass deliberately stopped before editing the scorecard, report, or sitews12-pace-sequencing.md records the evidence and a provisional row; PACE stays out of published rankings until the full board, whitepaper and site can move together
24, 26The PACE evidence pass began without PACE-specific M3 adjudication criteria, and the criteria were written afterwardsThe result is labelled exploratory and provisional; the criteria govern only future evidence, and cannot be used to upgrade the pass that preceded them
25, 27A CMS PACE directory enumeration was added after the initial evidence review, then corrected when an independent recheck found 201 visible enrollment values totalling 79,758 with five suppressed rowsReported as a visible-enrollment lower bound with the file's masking and coverage limits stated; does not upgrade the causal scalability verdict
28The prospective PACE evidence check produced no qualifying support or refutation for any of the four new M3 criteriaAll four criterion-level cells remain explicitly indeterminate — which is why PACE is still not on the published board
23, 29Completed the line-edit pass queued at #19 across six findings files; added a prospective HCBS-rebalancing criterion after the original workforce passUnderlying files now carry the corrected scopes and tiers (Phase 1 verification later found two lines the pass missed, since fixed). §13 records narrow package-level support for BIP-style rebalancing only
35–37Independent Phase 2 steelman against the filing's verdict, its top-ranked architecture and its scorecard headline (2026-08-10)Two published claims withdrawn — "the best-evidenced fix isn't the one anyone's proposing" (it has been federal law since 2005 and ~1.5M people use it) and the unqualified verdict "Medicare mostly doesn't pay for this". Also found that Pub. L. 119-21 — cited on the whitepaper for §71111 — partially enacts the HCBS de-capping architecture at §71121, and that two studies the filing cites each contain a second finding it did not carry across. See steelman-log
32–34Independent Phase 1 fact-check against primary statute, rule and paper text (2026-08-10)40 claims corrected and 42 reworded across these pages and the research record; full ledger and log published — see verification-log
30Ran the structurally blinded re-score this filing had been claiming since publication but never had — and amended the anchored scale in the course of it, after finding that high scores required cited evidence while low scores accepted its absenceThe blind matrix differed on 26 of 40 ranked cells; 25 were corrected and 1 kept as a logged judgment split, so the board changes on 25 of 40 (14 matched outright). One headline withdrawn as stated; one strengthened; the wage floor's workforce lead lost under both scale readings, so the filing's own sequencing prior draws no support from its own board. Because the scale amendment is load-bearing, both readings are published side by side
31Three passages leaking prior cell values were redacted from the blinded evidence packet before it was sent, and PACE was scored from an evidence-only extract of the sequencing fileEach redaction left a visible marker rather than a silent deletion, and all three are itemised in the raw-output file along with the residual risks that were deliberately not redacted. PACE and the Credit for Caring Act now have independent blind rows recorded; neither is entered on the published board
15Four workstreams (§2 spending, §4 insurance reform, §7 assisted living, §10 international) dispatched in one parallel batch rather than sequentially, at the user's request to pick up paceNo QA shortcuts taken — each report still independently reviewed before being written into a findings file; parallelizing changed wall-clock only
17§7's H7.2 (Medicaid AL footprint growth) and the AL-vs-SNF incident-rate comparison both rest on directional evidence, with no clean longitudinal or rate-comparison dataset foundRecorded as explicit data gaps rather than filled with a plausible-but-unverified inference
18§10's Japan findings rely partly on a methodological critique (Geyer) whose core claim was not independently verified against the critique's own primary dataRecorded as the critique's claim, attributed to its source, not adopted as this protocol's own finding
19A full red-team pass (11 agents) produced 22 resolved items and 3 reconciled scorecard cells, but corrections were applied to research-inquiry.md, red-team-log.md, ws11-scorecard.md, and ws05-findings.md — not line-edited into the prose of ws03, ws04, ws06-09, ws07, ws08, and ws10-findings.md, which can still read pre-red-team in placesA disclosed, not hidden, inconsistency: the whitepaper incorporates every correction, but a reader opening those six files directly without also reading red-team-log.md will see uncorrected verdict language. Line-edit pass queued
20The first whitepaper build used a stale, superseded design system (brand/gubment.css, childcare/report/ fonts) instead of the live site's actual current templateIncorrect build deleted, not left alongside the correct one; rebuilt and screenshot-checked against the live site. brand/ and childcare/report/ flagged for deprecation cleanup before GBMT-8/9
21The published whitepaper references an og-card.png that doesn't yet exist for elder care specifically — it will fall back to childcare's card if shared on social media before a real one is madeNot a broken link, but a wrong one until an elder-care-specific og-card.png is generated; queued as a later plain-talk-artifact stage per M8
Record 4

Red team

22 attacks on the workstream findings and scorecard, run as a Workflow (four attack/defend groups plus a re-score, reconciliation, and synthesis pass), each answered and, where it landed, absorbed into the record rather than defended against. This log is the authoritative correction layer — the whitepaper incorporates every item below, though the individual workstream files may still read pre-red-team in places (see deviations-log #19).

Attack 1 — the CLASS/WA Cares "cleanest natural experiment" claim

Partially lands. Benefit-design variance (WA Cares' capped $36,500 vs. CLASS's proposed uncapped benefit) was unaddressed — a design that fails on adverse selection and one that fails on an inadequate benefit could look identical without payout data, and there is none yet. Corrected to "the cleanest available contrast on the adverse-selection variable specifically, confounded by uncontrolled benefit-design differences."

Attack 2 — anchor 11's (AL vs. SNF) flat "prior broken" table verdict

Absorbed. The prose already disclosed the 3-year vintage gap and paywalled AL data; the table contradicted its own prose. Corrected to "prior likely broken on current data, but vintage-mismatched and directionally unstable — do not treat the ~20% margin as current."

Attack 3 — the caregiver-valuation "cannot be reconciled" framing

Absorbed. All three divergence drivers (scope, wage basis, hours source) are named and independently explain the spread — that's attribution, not irreconcilability. Corrected to "reconciled to three named methodological drivers, not collapsible to a single point estimate"; the instruction to report a conditional range stands unchanged.

Attack 4 — anchor 2's (PHI jobs) "undersells ~10x" framing

Absorbed. The 10x figure compared growth-only against growth+replacement, a category mismatch. Like-for-like is 772K vs. "1M+" — roughly 25% off. The 10x framing is retired.

Attack 5 — H3.1's "directionally suggestive" verdict

Absorbed. Economy-wide post-pandemic wage inflation across nearly every low-wage service sector is an unaddressed confound producing the identical wages-up/vacancies-flat pattern with zero HCBS-specific content. Verdict downgraded to plain INDETERMINATE, confidence low.

Attack 6 — §2 spending's "$415B, medium-high confidence" tier

Partially lands. KFF's multiple compilations share upstream NHEA accounting assumptions — restatements of one convention, not independent cross-checks. Tier stands, but the "multiple sources converge" framing needed that caveat made explicit.

Attack 7 — ws04's "reinsurance doesn't touch adverse selection" claim

Answered, no change. The "little potential to generate savings" finding is Milliman's own; the adverse-selection framing is clearly labeled as the researcher's bridging inference, not misattributed as Milliman's claim.

Attack 8 — ws04's "tested and rejected" framing

Absorbed. Action language outran a single unread, secondhand-reported state study. Softened to "suggestive negative evidence from one unread secondary-sourced state study."

Attack 9 — ws04's "reinsurance" generalized as a category

Absorbed. The finding applies specifically to Washington's state-subsidized reinsurance design; scope narrowed to "this reinsurance design" — a federal design with different mechanics remains untested.

Attack 10 — "BPC favoring auto-enrollment is itself a tell"

Absorbed. Confirmation-bias framing dropped; no competing federal reinsurance proposal was found, but that's equally explained by upfront cost asymmetry as by a verdict on mechanism efficacy.

Attack 11 — ws05's "SUPPORTED in Indiana" header

Absorbed. Body already disclosed the 2024 waiver-restructuring confound, but the bold header overstated it. Corrected to "CONSISTENT WITH H5.1 (n=1, correlational, confounded)."

Attack 12 — "no evidence of new state legislation reacting to rescission"

Absorbed. Six months (Feb–Aug 2026) is a thin window for most state legislatures. Corrected to "not yet visible given a ~6-month post-rescission window and reporting lag" — not read as states declining to respond.

Attack 13 — the Health Affairs closures-finding's scope

Absorbed. The 22-state panel studied mandates already in force (up to ~4.1 HPRD), not the rescinded federal 3.48 HPRD floor applied uniformly. Corrected to scope the finding to mandates studied, not ratio mandates in general.

Attack 14 — PE-mortality confidence rated "high, barely contested"

Partially lands. A JAMA COVID-era comparison found PE facilities' pandemic mortality in line with industry average, excluded by the study's own co-author as inconclusive rather than reconciled. Confidence downgraded from high to medium-high; the core +11% estimate remains unrebutted.

Attack 15 — the New York enforcement-gap generalization

Answered, no change. NY is likely the most-enforced/visible case, which if anything strengthens rather than weakens the finding's existing caution that a less-visible state could show worse enforcement.

Attack 16 — §7's H7.2 "directionally plausible" framing

Absorbed. The cited trend (HCBS spending passing institutional spending in 2013) is overwhelmingly about home-based care, not assisted living, and is equally consistent with AL holding a flat/shrinking share of a growing category. Corrected to "consistent with, but not evidence for, de facto institutionalization."

Attack 17 — ws11's merged Cash & Counseling / Credit for Caring row

Absorbed. ws08 keeps the two mechanisms separate throughout; the merge happened only in ws11's row label. Split into two rows — Cash & Counseling keeps its RCT-backed relief score; Credit for Caring demoted and left unscored on other objectives, since no empirical test of the bill itself exists.

Attack 18 — ws08's caregiver-health header "solid, federal-sourced, and getting worse"

Absorbed. Bundles a dated 2011 MetLife dollar figure with current CDC/AARP data. Corrected to "federal-sourced where it matters; the dollar-valuation figure remains a dated private estimate."

Attack 19 — ws08's BRFSS "widened, not narrowed" trend (2015–16 to 2021–22)

Partially lands. The endpoint sits inside the pandemic window; the trend hasn't been shown to persist post-recovery. Downgraded from "strongest non-advocacy evidence" to suggestive, confounded by pandemic timing.

Attack 20 — ws10's Japan "traces to one research lineage" claim

Partially lands. True that citations cluster around one author lineage, but no systematic search for independent replications was conducted. Kept the flag, dropped the implied completeness.

Attack 21 — ws10's cross-cutting strain claim

Answered, minor addition. Already correctly scoped to Germany/Japan/Netherlands inside the measurement window; added that no non-strained counterexample was sought.

Attack 22 — ws11's AL-standard "same vulnerability, arguably worse" framing

Absorbed. No Conditions-of-Participation framework exists yet for assisted living, so a court extending the SNF rule's exact reasoning is untested, not established. The load-bearing part is the political precedent, not a demonstrated legal one; "arguably worse" retired.

A fresh scorer, blind to the scorecard's original numbers, independently re-scored all nine architectures from the underlying finding files alone; three cells differed by ≥2 points and were reconciled: Cash & Counseling's cost-containment score corrected 3→1 (its own RCT documented a real year-one cost increase, not merely absent savings evidence); the federal assisted-living standard's workforce-sustainability score corrected 1→3 (the Health Affairs staffing-mandate evidence applies here too and was missed, not reasoned away); OAA/NFCSP's cost-containment score corrected 4→2 (the funding-scale comparison is not itself a cost-containment finding). Consequence, caught by the reconciliation itself: Cash & Counseling no longer wins under every weighting, only under caregiver-relief weighting where its RCT-backed score dominates.

Correction, 2026-08-10 — this re-score was weaker than we described it. The scorer above was an agent instructed not to open the scorecard while holding full file access throughout. That is instruction-level blinding, not the structural blinding our method requires, and this page previously presented it as the genuine article. A structurally blinded pass — one request, no tools, evidence inlined, the scorecard unreachable by any route — was run on 2026-08-10 and differed on 26 of the 40 ranked cells, moving 25 (one difference was kept at its published value as a logged judgment split). This paragraph previously said “moved 26”; corrected on Phase 1 verification, 2026-08-10, to match every other statement of the same fact. All three corrections above were independently re-derived by it, so they stand. The sentence that used to close this paragraph — "Private LTC insurance market reform still ranks last under every weighting tested, unchanged" — does not, and is withdrawn: four of that row's five cells rested on no evidence in either direction, and it ranks last only under the original, one-directional scale. Both readings are published side by side in ws11-rescore-log.md.

Record 5

Phase 2 steelman — the strongest case against our three loudest claims

Run 2026-08-10 in a separate session from the Phase 1 fact-check, per the Verification Protocol's independence rule. The targets were derived from this filing's own text before any evidence was gathered: the verdict printed on every screen, the architecture that leads the board, and the headline over the scorecard. Findings that weaken the filing are the point of this pass; two of the three targets broke.

Target A — "Medicare mostly doesn't pay for this"

Partially survives; verdict narrowed. Right about custodial care, and right at statutory tier: 42 U.S.C. §1395y(a)(9) excludes it by name. Wrong as a statement about paid elder care. In CMS's own National Health Expenditure Accounts for 2024, Medicare is the largest single payer of home health care in the country (32.9%, ahead of Medicaid) and pays 21.5% of nursing-facility care. Since plan year 2020, Medicare Advantage plans — covering 54% of beneficiaries — may buy supplemental benefits that "may not be limited to being primarily health related" for chronically ill enrollees. And CMS's GUIDE model has paid up to $2,500 a year in caregiver respite, nationwide, since July 2024. None of that appears anywhere in this filing: "Medicare Advantage" returned zero hits in the whole record. What defeats the steelman's stronger form: MedPAC states outright that Medicare's data cannot measure supplemental-benefit use, so the size of that channel is unknown.

Target B — HCBS de-capping, the architecture that leads our board

Survives the objection; the framing does not. We built the standard woodwork case against expanding home care and it lost — on our own citation. The study we cite for "$0.74" is titled Is there a "woodwork" effect? and answers no, reporting that dollar as an offset and finding more older adults served at lower cost per beneficiary. We had published only the qualifying half. But two things we said are wrong in shape. The waiver cap is not only institutional bias: 42 U.S.C. §1396n(c)(2)(D) makes per-person cost-neutrality against institutional care the condition of the waiver, and an enrollment cap is how a state meets it. And Pub. L. 119-21 §71121 — the same statute we cite on the whitepaper for killing the staffing rule — partially enacts this architecture, letting the Secretary approve waivers from July 1, 2028 for people who do not meet the institutional level-of-care test, with $150M of implementation money. The enacted version keeps the cost cap, requires that existing waiting lists not lengthen materially, and bars the money from funding direct-care workers' health insurance or training.

Target C — "the best-evidenced fix isn't the one anyone's proposing"

Steelman wins; claim withdrawn. Congress enacted the Cash and Counseling design as a permanent state-plan option in 2005 (42 U.S.C. §1396n(j)) and again in the ACA as Community First Choice (§1396n(k)), attaching a six-percentage-point federal match bonus. Self-direction runs in all fifty states and DC; about 1.5 million people use it, up 23% since 2019. Our own ws08-findings.md said so and this page said the opposite. Sharper still: the paper supplying this board's strongest positive cell evaluates two programs and reports no significant effect for Community First Choice (1.51%, CI −12.77% to 15.79%) beside the Balancing Incentive Program's +13.24%. We cited the arm that helped. The underlying finding survives — Cash and Counseling is still the only architecture here anchored by a randomized trial — but it is not an unproposed fix; it is a scaled one nobody evaluated.

The pattern, which is the useful part. Three of this pass's load-bearing findings are second halves of documents we already cite — a null in the same abstract, a title that answers its own question, a section of a statute we cite by number. Nothing was suppressed. Each time, the half that fit our frame was published and the half that complicated it was not. Phase 1 found the same shape from the other direction. Full case, including the two points where the steelman lost outright and the scorecard sensitivity it produced: steelman-log.md.

THE RECEIPTS · 11-workstream protocol (9 executed to findings) · pre-registered anchor table (M4) preserving corrected priors · committed ACS population-baseline pipeline · deviations log (37 entries) · full red team applied (22 items, 3 reconciled scorecard cells) · a second, structurally blinded re-score (25 of 40 cells moved, one headline withdrawn, both scale readings published) · an independent Phase 1 fact-check against primary sources (689 claims, 40 corrected, 42 reworded) · an independent Phase 2 steelman (two published claims withdrawn) · all sourced above, full detail in the repository.