GUBMENTPlain talk · policy frontier
Filings / Elder care / Sources / Red Team Log (GBMT-7)
GBMT-7 · Research record · No. 7

Red Team Log (GBMT-7)

elder-care/research/red-team-log.md
This is a working research document from the elder care filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 7, not a summary of it.

Per M6: written attacks on the draft's conclusions, each answered or absorbed, confidence downgrades applied to the record. Run as a Workflow (4 groups × attack/defend, dispatched in parallel) against every workstream finding file plus an independent blind re-score of the scorecard. This log is the authoritative correction layer — whitepaper-07-content.md incorporates every item below; the individual ws0X-findings.md files may still read pre-red-team in prose (see deviations-log #19).

Group: Phase 0 + §2 baseline + §3 workforce

  1. CLASS/WA Cares "cleanest natural experiment" — PARTIAL. Benefit-design variance (WA Cares' capped $36,500 vs. CLASS's proposed uncapped ongoing cash) is real and was unaddressed. A design that fails on adverse selection and one that fails on an inadequate/capped benefit could look identical from outside absent payout data — and there is none yet. Corrected: "the cleanest available contrast on the adverse-selection variable specifically, confounded by uncontrolled benefit-design differences" — not "clean" or "controlled" unqualified.
  2. Anchor 11 (AL vs. SNF) — ABSORBED. The prose already disclosed the 3-year vintage gap and paywalled AL data; the table's flat "Prior broken... ~20%" contradicted its own prose. Corrected: "Prior likely broken on current data, but vintage-mismatched (2022 vs. 2025) and directionally unstable — do not treat the ~20% margin as current." Applied to research-inquiry.md row 11.
  3. Caregiver-valuation "cannot be reconciled" — ABSORBED. All three divergence drivers (scope, wage basis, hours source) are named and independently explain the spread — that's attribution, not irreconcilability. Corrected: "reconciled to three named methodological drivers, not collapsible to a single point estimate." The downstream instruction (report a conditional range, never one number) is unchanged.
  4. Anchor 2 (PHI jobs "undersells ~10x") — ABSORBED. The 10x figure compared growth-only against growth+replacement, a category mismatch. Like-for-like is 772K vs. "1M+" — roughly 25% off. Corrected: retire the 10x framing; anchor delta now reads "~25% undersold on the like-for-like growth figure." Applied to research-inquiry.md row 2.
  5. H3.1 "directionally suggestive" — ABSORBED. Economy-wide post-pandemic wage inflation (2021–2024, nearly every low-wage service sector) is an unaddressed confound that produces the identical wages-up/vacancies-flat pattern with zero HCBS-specific content — no matched comparison sector was checked. Corrected: verdict downgraded to plain INDETERMINATE; confidence tier low, not low-to-moderate, until a macro-confound control is attempted.
  6. §2 spending, $415B medium-high confidence — PARTIAL. KFF's multiple compilations share upstream NHEA accounting assumptions — they're restatements of one convention, not independent cross-checks of each other. Tier stands, but the "multiple sources converge" framing needed this caveat made explicit rather than implied.

Group: §4 financing + §5 HCBS access

  1. ws04 "reinsurance doesn't touch adverse selection" — ANSWERED. The doc's "little potential to generate savings" is Milliman's actual finding; "doesn't touch the adverse-selection mechanism" is clearly labeled as the researcher's bridging inference to the CLASS diagnosis, not misattributed as Milliman's own claim. No change; prose tightened for clarity only.
  2. ws04 "tested and rejected" — ABSORBED. Action language outran a single unread, secondhand-reported state study. Corrected: CR score and framing softened to "suggestive negative evidence from one unread secondary-sourced state study," not a definitive rejection.
  3. ws04 "reinsurance" generalized as a category — ABSORBED. The finding applies to Washington's specific state-subsidized reinsurance design (state reimburses private insurers for catastrophic claims, voluntary market). Corrected: scope every reference to "this reinsurance design," not "reinsurance" broadly — a federal design with different risk-corridor/subsidy/mandate mechanics is untested.
  4. ws04 "BPC favoring auto-enrollment is itself a tell" — ABSORBED. Confirmation-bias framing. Corrected: neutral note — no competing federal reinsurance proposal was found, but that's equally explained by political/fiscal cost asymmetry (reinsurance needs upfront state spending, auto-enrollment doesn't) as by a verdict on mechanism efficacy. Drop "is itself a tell."
  5. ws05 "SUPPORTED in Indiana" header — ABSORBED. Body already disclosed the confound (2024 waiver-restructuring timing) but the bold header overstated it. Corrected header: "CONSISTENT WITH H5.1 (n=1, correlational, confounded)." Applied to ws05-findings.md.

Group: §6/§9 institutional quality/federalism + §7 assisted living

  1. §6/§9 "no evidence of new state legislation reacting to rescission" — ABSORBED. Six months (Feb–Aug 2026) is a thin window for most state legislatures. Corrected: "not yet visible given a ~6-month post-rescission window and legislative reporting lag" — not read as states declining to respond.
  2. §6/§9 Health Affairs closures finding scope — ABSORBED. The 22-state panel studied mandates already in force (up to ~4.1 HPRD, DC), not the rescinded federal 3.48 HPRD floor applied uniformly. Corrected: "undercuts the closures argument for the range of mandates studied (up to ~4.1 HPRD)" — not ratio mandates in general, and doesn't speak to whether a stricter uniform national floor crosses a different threshold.
  3. §6/§9 PE-mortality "high, barely contested" — PARTIAL. A JAMA COVID-era comparison found PE facilities' pandemic mortality in line with industry average; the study's own co-author acknowledged excluding it as inconclusive rather than reconciling it. Corrected: confidence downgraded from high to medium-high — the core IV estimate (+11%) remains methodologically strong and econometrically unrebutted; this is a tightening, not a reversal.
  4. §6/§9 NY enforcement-gap generalization — ANSWERED, no change. NY is likely the most-enforced/visible case (litigation + media coverage), which if anything strengthens rather than weakens the finding's existing caution that a less-visible state could show worse, not better, enforcement.
  5. §7 H7.2 "directionally plausible" — ABSORBED. The cited trend (national HCBS spending passing institutional spending in 2013) is overwhelmingly about home-based care, not assisted living, and is equally consistent with AL holding a flat/shrinking share of a growing category. Corrected: "consistent with, but not evidence for, de facto institutionalization" — the gap isn't just imprecision, the cited trend doesn't discriminate between the hypothesis and its negation.

Group: §8 family caregiving + §10 international + §11 scorecard

  1. ws11 "Cash & Counseling / Credit for Caring" merged row (CR=5) — ABSORBED. ws08 keeps the two mechanisms separate throughout (different eligible populations, no shared administrative machinery, no CBO score for the credit); the merge happened only in ws11's row label. Corrected: split into two rows — Cash & Counseling keeps CR=5 (RCT-backed); Credit for Caring Act demoted to CR 2–3, unscored on other objectives (no empirical test exists for the bill itself).
  2. ws08 caregiver-health header "solid, federal-sourced, and getting worse" — ABSORBED. Bundles the 2011 MetLife figure (dated, private, no successor study) with the CDC/AARP data (federal, current). Corrected: "federal-sourced where it matters (BRFSS, AARP labor data); the dollar-valuation figure remains a dated private estimate."
  3. ws08 BRFSS "widened, not narrowed" (2015–16 to 2021–22) — PARTIAL. The endpoint sits inside the pandemic window; the trend hasn't been shown to persist post-recovery. Corrected: downgrade from "strongest non-advocacy evidence" to suggestive, confounded by pandemic timing, until a post-2022 wave is checked.
  4. ws10 Japan "traces substantially to one research lineage" — PARTIAL. True that citations cluster around Campbell/Ikegami/Gibson, but no systematic search for independent replications was conducted. Corrected: "the favorable narrative's most-cited sources cluster around one research lineage; no systematic search for independent replications was conducted" — keep the flag, drop the implied completeness.
  5. ws10 cross-cutting strain claim — ANSWERED, minor addition. Already correctly scoped to Germany/Japan/Netherlands inside this window, not generalized. Added: no non-strained counterexample was sought.
  6. ws11 AL-standard "same vulnerability, arguably worse" — ABSORBED. No Conditions-of-Participation framework exists yet for AL, so a court extending the SNF rule's exact reasoning is untested, not established. Corrected: the load-bearing part of this row is the political precedent (a recent bipartisan federal reversal of an analogous rule), not a demonstrated legal one; "arguably worse" retired.

Independent re-score and reconciliation

Superseded 2026-08-10 — read this section with the correction below. This re-score was an in-workflow agent instructed not to open ws11-scorecard.md while holding filesystem access throughout: instruction-level blinding, not the structural blinding M6 requires. A structurally blinded pass (ws11-rescore-log.md) ran on 2026-08-10 and moved 25 of the 40 ranked cells. All three corrections below were independently re-derived by it and stand. The consequence paragraph that closes this section does not — see the note appended to it.

A fresh scorer, blind to ws11-scorecard.md's original numbers, re-scored all 9 architectures from the underlying finding files alone. Three cells differed by ≥2 points and were reconciled against the source files directly:

Cell Original Rescore Reconciled Why
Cash & Counseling — CC 3 1 1 The RCT's own year-one cost data (AR +17%, FL +14%, both significant) is direct evidence of a cost increase, not merely absent evidence of savings — the anchored scale's 1 is defined for exactly this case.
Federal AL standard — WS 1 3 3 The Health Affairs 22-state panel (already used to justify WS=3 on the wage-floor row) applies here too and was missed, not reasoned away — a genuine partial, state-level positive analog.
OAA/NFCSP — CC 4 2 2 The 209M − vs140B comparison is a scale finding, not a cost-containment finding — no file evaluates whether NFCSP funding contains system-wide costs. Original conflated implementation-affordability with the cost-containment objective as anchored.

Consequence, caught by the synthesis pass itself, not smoothed over: the reconciled Cash & Counseling CC score means it no longer wins under every weighting. Recomputed: it now ties or trails nationalized WA Cares insurance and the wage floor under cost-containment and workforce weighting, and leads clearly only under caregiver-relief weighting. Private LTC insurance market reform still ranks last under all three weightings, no exceptions. Full corrected scorecard and rank-stability discussion: ws11-scorecard.md.

Correction, 2026-08-10. Two claims in the paragraph above are withdrawn by the structurally blinded re-score. "Private LTC insurance market reform still ranks last under all three weightings, no exceptions" — four of that row's five cells rested on no record evidence in either direction; it places seventh of eight under the amended scale and last only under the original one-directional scale, which makes it the least stable finding on the board rather than the most. "It now ties or trails nationalized WA Cares insurance and the wage floor" (of Cash & Counseling) — it beats nationalized WA Cares insurance under all three weightings and the wage floor under two of three. What survives, and is the half that matters: Cash & Counseling does not win under every weighting. Both readings published in ws11-rescore-log.md.

Sources

Workflow wf_880d6e84-404 (11 agents: 4 attack + 4 defend + 1 independent score + 1 reconciliation + 1 synthesis), run 2026-08-03.

← All Elder care research documents Sources digest Read the whitepaper