GUBMENTPlain talk · policy frontier
Filings / Elder care / Sources / Blind re-score — raw scorer output (GB
GBMT-7 · Research record · No. 7

Blind re-score — raw scorer output (GBMT-7)

elder-care/research/ws11-blind-scores-2026-08-10.md
This is a working research document from the elder care filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 7, not a summary of it.

Date: 2026-08-10 Model: claude-opus-5 via the Anthropic Message Batches API (scripts/batch-rescore.py, batch msgbatch_01U9X84UsaJNvZ8YtrMFzaDZ). Blinding: structural. The scorer received one tool-less request containing the scoring brief, the anchored scale, the scoring disciplines, and twelve evidence documents inlined as text. It had no filesystem access and no search, so it could not reach ws11-scorecard.md, red-team-log.md, or deviations-log.md by any route. This is stronger than the 2026-08-03 pass-2 re-score, which was an in-workflow agent instructed not to look.

What the scorer received. docs/phase0-findings.md, docs/research-inquiry.md, and research/ files ws02, ws03, ws04, ws05, ws06-09, ws07, ws08, ws10, ws13, plus an evidence-only extract of ws12.

What it did not receive. ws11-scorecard.md, red-team-log.md, deviations-log.md, site/elder-care/*, elder-care/report/*, and ws12's provisional scorecard row, sequencing recommendation, and M3 adjudication tables.

Redactions made to preserve blinding, disclosed in full. Three passages in the evidence base leaked prior scorecard values or named a target cell, and were replaced with a visible [REDACTED FROM THIS PACKET: …] marker rather than silently deleted:

  1. ws04-insurance-reform-findings.md — the sentence beginning "Update ws11-scorecard.md's private-LTC-insurance-reform row", which states four of that row's five cell values verbatim.
  2. ws04-insurance-reform-findings.md — the opening line "Resolves the row ws11-scorecard.md flagged as untested."
  3. ws12-pace-sequencing.md — the clause "the honest cell is weak/negative rather than 'unmeasured'", which names the cell the CC score should take.

Residual contamination risk, disclosed rather than asserted away. The packet retains dangling markdown links to red-team-log.md and deviations-log.md (unreachable — the scorer had no tools) and retains ws13's "Scorecard boundary" paragraph, which instructs that the HCBS row not be re-scored without an independent pass and that the partial-offset cost caveat be preserved. That instruction was kept deliberately: scoring discipline 6 requires honouring an instruction not to resolve a tradeoff, and removing it would have removed a legitimate constraint rather than a leak. ws13 also characterises the HCBS row's evidence as "stronger … than an ARPA-only record", which is a directional hint; it is left in for the same reason and flagged here.

One material change to the instrument, flagged because it drives most of the divergence below. The brief's anchor 3 was widened from the published scale's "mixed/indeterminate — plausible mechanism, but real countervailing evidence" to also cover "no evidence in the record either way", and anchor 1 was narrowed to require documented adverse effect rather than a mechanism that "plausibly cuts the wrong way". This imports the house evidence-floor discipline applied to childcare, housing, media and drugs in 2026-08. The reasoning, and the case that the published scale is structurally defective, is in ws11-rescore-log.md.

Reconciled cell by cell in ws11-rescore-log.md. What follows is the scorer's output, verbatim and unedited.


1. The matrix

CC IQ WS CR AA
A1 2 3 2 3 3
A2 2 3 4 3 4
A3 3 3 3 3 3
A4 1 3 3 5 5
A5 3 3 3 2 3
A6 2 3 3 3 4
A7 3 3 4 3 3
A8 3 3 3 3 3
A9 2 3 3 3 3
A10 3 3 3 3 3

2. Cell-by-cell basis

A1 — Federal LTC social insurance (WA Cares nationalized / CLASS / DE-JP-NL)

Cell Score Evidence Flag
A1CC 2 10-ws10-international.md: Germany's contribution rate rose again for 2025 and federal loans were needed to hold 2026 rates steady; Netherlands' Wlz is the highest LTC spend in the OECD (4.4% GDP → projected 6% by 2070) and the 2015 reform's own cost-containment goal is partly undermined by perverse incentives; Japan is proposing copay increases toward 30% — "active benefit-tightening, not a stable equilibrium." Plus 00-phase0-findings.md §4: HHS's 19-month CLASS actuarial review could not price the design into solvency under any fix tried. Real cited findings undermining cost containment; WA Cares' capped $36,500 benefit is a cost-control design feature with zero outcome data (first payout July 2026), which I treated as mechanism, not evidence. evidenced
A1IQ 3 Nothing in the record links a payroll-tax LTC benefit to institutional quality or safety outcomes in any of the five named instances. unevidenced-neutral
A1WS 2 10-ws10-international.md: under Pflegeversicherung, German direct-care vacancies rose 110% in a decade against 45% demand growth and 13% workforce growth (NBER/Geyer); Japan's own ministry projects a 570,000-worker shortfall by 2040; Netherlands projects 266,000 short by 2035 (44,000 already). The workstream's own framing: "Germany faces its own version of §3's binding constraint, not a solved problem to import." Judgment call: this documents shortage coexisting with the instrument over the measurement window, not a causal harm from it; a reasonable scorer could hold this at 3 evidenced-wash. I scored 2 because the record explicitly presents it as a finding against importing these systems as workforce solutions. evidenced / judgment-call
A1CR 3 Germany/NL/WA Cares permit cash benefits (noted only in contrast to Japan's in-kind rigidity in 10-ws10), but no caregiver-burden outcome for any of them appears in the record. WA Cares has ~0 months of payout data (00-phase0 §6, anchor 6). unevidenced-neutral
A1AA 3 No positive evidence. Flagged near-miss: 10-ws10 reports that the Dutch Wlz split gives municipalities and insurers an incentive to push people into the nationally-funded (severe/institutional) tier rather than cheaper municipal home care — the closest thing in the record to adverse AA evidence for a member of this instrument class. I did not let it drive the score because it is framed as a cost-containment finding about a two-tier federal/local split, with no measured institutionalization outcome, and does not describe the WA Cares/CLASS design this row is anchored on. Disclosed as a judgment call. unevidenced-neutral / judgment-call

A2 — Medicaid HCBS waiver expansion / de-capping

Cell Score Evidence Flag
A2CC 2 13-ws13-rebalancing.md: a 2023 DID found each additional dollar of HCBS spending associated with $0.74 of added total LTSS spending (only partial institutional offset) — "the self-financing claim is therefore not supported"; the 2026 IV/border study's 1.02pp reduction in Medicare community home-health use is "relevant cross-payer substitution, not evidence of total Medicare-plus-Medicaid savings." ws13 instructs preserving the partial-offset caveat, which I have: partial offset exists, but net cost rises. evidenced
A2IQ 3 Nothing in the record bears on the quality/safety of care inside institutional or residential settings under HCBS expansion. Reduced long nursing-home stays (below) is diversion, not institutional quality. unevidenced-neutral
A2WS 4 13-ws13-rebalancing.md: 2026 Health Services Research stacked DID (ACS 2005–21) found BIP adoption increased HCBS workforce by 13.24%, with no significant institutional-workforce effect (i.e. not mere shifting). Held at 4 rather than 5 because ws13 explicitly states this supports "the bundled, target-bound BIP design, not a generic claim that more HCBS money automatically creates staff," and there is countervailing evidence: a 2025 state-year study found no significant wage association with HCBS spending, and ARPA §9817's workforce-directed $26.3B left vacancy/turnover unimproved (03-ws03). Judgment call: 4 vs 5. evidenced / judgment-call
A2CR 3 No evidence in the record on HCBS expansion's effect on unpaid family caregivers. unevidenced-neutral
A2AA 4 13-ws13-rebalancing.md: 2025 BMJ Open generalized DID found BIP completer states had ~5 percentage points fewer long nursing-home stays for Medicaid beneficiaries living alone; the 2026 IV/border study measures real HCBS substitution. Held at 4, not 5, on the study's own named limits (HRS does not observe HCBS use directly; concurrent MLTSS adoption is a confound) and against the countervailing national picture — waitlists grew from ~692,679 (2021) to ~710,000+ (2024) through the ARPA window (03-ws03), though only 5% of §9817 money targeted waitlists and ws03 calls that pattern indeterminate. evidenced / judgment-call

A3 — Direct-care wage floor tied to funding

Cell Score Evidence Flag
A3CC 3 No source in the record measures total public or system LTSS cost under a wage floor or pass-through requirement. 06-ws06-09's Health Affairs panel finding (labor costs rose less than revenue, margins statistically unchanged) is about facility finances, not payer cost. I deliberately did not score down on the mechanism argument that a wage floor is cost-increasing by construction. unevidenced-neutral
A3IQ 3 No evidence links wage pass-through to institutional quality or safety outcomes. The Health Affairs +5% staffing result is an effect of staffing mandates (scored at A7), not of wage pass-through; the closest pass-through instrument in the record — NY's 2021 law directing 40% of revenue to resident-facing staffing — is documented only as an enforcement failure (~400 violating facilities, essentially no penalties for years), which is evidence about implementation, not about quality outcomes. unevidenced-neutral
A3WS 3 Evidenced wash inside a single objective. Positive: ARPA §9817 put 71% ($26.3B) into workforce recruitment/retention and "wages moved substantially where states used the money that way" — Colorado 12.41→ 18; national median home-care wage $13.07 (2014, infl-adj) → $16.77 (2024) (03-ws03). Negative/null: every responding state in KFF's 2023 50-state survey still reported shortages despite near-universal rate increases; Pennsylvania turnover held at 44–65%; a 2025 state-year study found no significant wage association with HCBS spending (13-ws13). ws03 explicitly instructs: "Call this INDETERMINATE, not SUPPORTED or directionally suggestive" (unresolved macro-wage confound, no §9817 evaluation requirement) — honoured per discipline 6. If wages alone counted, this would be a 4. evidenced-wash
A3CR 3 No evidence on family-caregiver outcomes from wage floors. unevidenced-neutral
A3AA 3 Evidence exists and fails to resolve: HCBS money spent overwhelmingly on wages coincided with waiting lists growing nationally (03-ws03), which cuts against community access — but ws03 attributes this to an untestable macro confound and notes only 5% of the money targeted waitlists, and Texas's ~95% duplicate-listing inflation (00-phase0 §5) makes waitlist counts a poor outcome measure. Not resolved in either direction. evidenced-wash

A4 — Cash & Counseling / self-direction expansion

Cell Score Evidence Flag
A4CC 1 08-ws08-family-caregiving.md: in the 1998–2003 RCT, cost was higher under treatment in year 1 in most states (AR elderly +17%, FL nonelderly +14%, both significant), narrowing to non-significant in Arkansas by year 2 but persisting in FL/NJ; the mechanism was closing unmet need — "it didn't do the same job cheaper, it did more of the job." This is direct experimental evidence of an adverse effect on this objective, which the scale's anchor 1 expressly covers. Judgment call, 1 vs 2: the AR year-2 convergence and the benign explanation (authorized-but-undelivered services) would support 2; I scored 1 because the adverse finding is RCT-grade and persisted in two of three states. evidenced / judgment-call
A4IQ 3 Nothing in the record on self-direction's effect on institutional or residential care quality — it is a community instrument and was not evaluated in those settings. unevidenced-neutral
A4WS 3 Explicitly an open gap, twice: 08-ws08 — whether self-direction substitutes for or complements the paid agency workforce is "explicitly flagged as untested by the researchers who'd know" (PMC12605748 calls it a hypothesis); 13-ws13 — "do not infer that paid-family-caregiver policy expands the agency workforce." Honoured per discipline 6 rather than resolved. unevidenced-neutral
A4CR 5 08-ws08: RCT (AR/FL/NJ, Mathematica/ASPE) — treatment caregivers reported less physical strain and better self-rated health across all three states. Confidence tier: high (primary ASPE/Mathematica reports). This is the only RCT-grade caregiver-relief finding in the entire record. evidenced
A4AA 5 08-ws08: RCT — satisfaction "consistently and substantially higher" under treatment (AR 60.8% vs 42.7% "very satisfied"), and the core finding that self-direction closed the gap where control-group members did not receive services they were authorized for; beneficiaries hire their own workers. Direct experimental evidence on the autonomy/control half of this objective. Flag: the RCT as reported does not measure institutionalization or community tenure, so the "keep people in the community" half rests on nothing; a scorer weighting that half more heavily could land at 4. evidenced / judgment-call

A5 — Credit for Caring Act

Cell Score Evidence Flag
A5CC 3 08-ws08: no CBO/JCT score exists, no independent think-tank static estimate was found; the workstream's own scale extrapolation ("bigger than NFCSP, smaller than HCBS") is stated as extrapolation, not a sourced number. No evidence either way on cost containment. unevidenced-neutral
A5IQ 3 No evidence. unevidenced-neutral
A5WS 3 No evidence. unevidenced-neutral
A5CR 2 08-ws08: S.925 / H.R.2036 (March 2025) is "the same bill, same fate as every prior version since 2016: referred to committee, no markup or floor action in any Congress it's been introduced in," and no fiscal or effect estimate of any kind exists; advocacy materials cite the size of the caregiving population, not an effect. The brief instructs scoring the bill "as drafted and as it has actually fared," and forbids importing Cash & Counseling's evidence. Ten years of non-advancement plus zero empirical support = anchor 2 ("no empirical support found"). Judgment call, disclosed: this is scoring political fate as delivery failure; if you read the brief as asking only about the design's effect-if-enacted, this cell is an unevidenced 3. judgment-call
A5AA 3 No evidence. unevidenced-neutral

A6 — PACE expansion

Cell Score Evidence Flag
A6CC 2 12-ws12-pace-evidence.md: the ASPE/Mathematica matched evaluation (8 states, 2006–2011) found Medicare spending mostly similar to predicted FFS while actual Medicaid capitation exceeded predicted Medicaid spending in every reported interval; ASPE's literature review reaches the same bottom line — no demonstrated savings and higher overall cost through Medicaid expenditure. ws12's verdict: "the record does not support a claim that PACE saves public money." Held at 2 rather than 1 because the two sources are not independent, the effect "varied materially by state," and ws12 rates confidence only "moderate on the direction of the cost-containment caution." Torn between 1 and 2. evidenced / judgment-call
A6IQ 3 Both directions present. Favorable: CMS/Abt found lower hospital admissions and fewer nursing-home days in the first six months. Against: differences narrowed at later follow-up, the report itself flags selection bias and small late samples, and ws12 states there is "no basis for calling PACE a universally proven quality/safety intervention." PACE is also community-based, so the evidence does not speak to quality in institutional/residential settings, which is what this objective asks. evidenced-wash
A6WS 3 12-ws12: workforce effects are explicitly not measured ("low on scalability and caregiver/workforce effects… those two latter effects are not measured by the material located here"). The eleven-role interdisciplinary-team requirement is stated as "a design fact, not a claim that PACE cannot scale," and ws12 retires "historically hard to scale" as NOT MEASURED. Honoured rather than converted into a mechanism-based downgrade. unevidenced-neutral
A6CR 3 12-ws12 lists "a direct measure of caregiver burden or labor-force outcomes attributable to PACE" among the things that would change the finding — i.e. it does not exist in the record. unevidenced-neutral
A6AA 4 12-ws12: CMS/Abt found fewer nursing-home days and more days in the community, plus initially higher life satisfaction; the 2024 JAMA Health Forum systematic review found PACE associated with reduced long-term nursing-home stays in three of four studies. Explicitly non-causal (declined-PACE comparison group, selection bias), so 4, not 5. evidenced

A7 — Federal minimum standards for assisted living

Cell Score Evidence Flag
A7CC 3 No evidence on payer/system cost from setting standards. The nearest datum — labor costs rising less than revenue with margins unchanged under state mandates (06-ws06-09) — concerns facility finances, not cost containment, and implies revenue (i.e. payments) rose. Not scored in either direction. unevidenced-neutral
A7IQ 3 Genuinely torn; real evidence both ways. For: 06-ws06-09 2025 Health Affairs 22-state panel (2010–2023) — mandates raised direct-care staffing ~5% with no effect on closure rates; 07-ws07 GAO-18-179 and GAO-26-107884 document a persistent, government-verified incident-reporting gap (only 22 of 48 states could report critical incidents; only 34 of 48 published any), with a real CMS reporting fix finalized May 2024 but not operative until 2027 at earliest. Against: the nearest federal analogue was killed twice over — vacated in N.D. Tex. and N.D. Iowa, HHS withdrew its appeals, Congress imposed a ~10-year moratorium via H.R. 1, CMS rescinded effective Feb 2, 2026, generating essentially no compliance record (00-phase0 §2); and where a pass-through-style state law exists, enforcement lagged badly (NY: ~400 violating facilities, essentially no penalties for years, fines only recently reaching ~20 facilities). Also, no state-by-state AL-vs-SNF per-resident incident-rate comparison exists, so the "AL is undercounted, not safer" inference is "plausible, not itself measured" (07-ws07). Nets out. evidenced-wash / judgment-call
A7WS 4 06-ws06-09: 2025 Health Affairs 22-state panel found staffing mandates raised direct-care staffing ~5% (CNAs +5.7%, LPNs +7.5%), labor costs rising less than revenue, margins statistically unchanged, and no effect on facility closures. That is a real workforce-headcount effect from a standards instrument, but in the SNF setting rather than AL and for mandates up to ~4.1 HPRD — anchor 4 (analogous natural experiment), not 5. Countervailing but not disqualifying: NY's enforcement gap, and the paper's own caution that a stricter uniform national floor need not behave the same. evidenced
A7CR 3 No evidence on family-caregiver outcomes from AL or SNF standards. unevidenced-neutral
A7AA 3 No evidence. Residential-care standards were not evaluated against community tenure or resident autonomy anywhere in the record. unevidenced-neutral

A8 — OAA / NFCSP expansion

Cell Score Evidence Flag
A8CC 3 No outcome evidence of any kind on NFCSP or Title III-C in the record — only the FY2026 appropriation (~$209M) and its ratio to HCBS spending. unevidenced-neutral
A8IQ 3 No evidence. unevidenced-neutral
A8WS 3 No evidence. unevidenced-neutral
A8CR 3 The record confirms NFCSP's scale only: ~$209M FY2026, "three orders of magnitude below Medicaid HCBS," funding "well under 0.1% of even the lowest independent valuation" of caregiver labor (01-protocol anchor 12, confirmed in 08-ws08). No evaluation of respite, counselling, or support-service outcomes appears anywhere in the packet. Explicit discipline note: the tiny ratio tempts a 2, but discipline 2 states that "scale descriptors, and the size of the problem being targeted are not evidence of an instrument's effect." Held at exactly 3. unevidenced-neutral
A8AA 3 No evidence. Home-delivered meals supporting community tenure is a mechanism argument only. unevidenced-neutral

A9 — Private LTC insurance reform (WA state-subsidized reinsurance)

Cell Score Evidence Flag
A9CC 2 04-ws04-insurance-reform.md: Washington's 2015 legislature-mandated Milliman feasibility study found the public-private reinsurance option had "little potential to generate savings"; the workstream labels this "real but scope-limited negative evidence, not 'tested and rejected,'" and notes the Milliman PDF was read only through secondary reporting. Held at 2, not 1: it is a null-on-savings for one state's design, secondary-sourced, and the workstream explicitly reserves judgment on federal risk-corridor/subsidy/mandate variants. evidenced
A9IQ 3 No evidence. unevidenced-neutral
A9WS 3 No evidence. unevidenced-neutral
A9CR 3 No evidence on family-caregiver outcomes. The record's adjacent findings (~7–10 standalone carriers left; hybrid products pitched at 500K5M net worth buyers, sidestepping rather than solving adverse selection; ws04's own note that shifting tail risk "does not, by itself, change who chooses to enroll" is the researcher's inference, "not a finding Milliman established") speak to market viability, not to caregiver burden, and ws04 explicitly labels the enrollment-mechanics inference as unestablished. unevidenced-neutral
A9AA 3 No evidence on where privately insured recipients receive care. unevidenced-neutral

A10 — Cash comparator (unconditional, unstructured)

Cell Score Evidence Flag
A10CC 3 The record contains no evaluation of unconditional cash transfers to older adults or families. unevidenced-neutral
A10IQ 3 No evidence. unevidenced-neutral
A10WS 3 No evidence. unevidenced-neutral
A10CR 3 No evidence. Explicit decision: I did not import the Cash & Counseling RCT here. A4 is C&C; A10 is defined as cash with no counselling, no fiscal intermediary and no oversight, and the C&C evaluation in the record does not decompose the cash component from the counselling/fiscal-intermediary components. Attributing the RCT to both rows would manufacture identical scores on evidence that cannot distinguish them — but so does declining to, in the opposite direction. Disclosed as a judgment call. unevidenced-neutral / judgment-call
A10AA 3 No evidence, same reasoning as A10CR. unevidenced-neutral

3. What your matrix implies

On my numbers A4 (Cash & Counseling / self-direction) is the strongest row and simultaneously carries the single worst cell on the board: it is the only architecture with RCT-grade positive evidence (5 on caregiver relief, 5 on autonomy) and the only one with RCT-grade adverse evidence on cost (1) — the same study produces both, because it cost more precisely by delivering care that the control group was authorized for and never received. A2 (HCBS de-capping/rebalancing) is the next strongest and the best-rounded: quasi-experimental positives on workforce (+13.24% HCBS workforce under BIP) and on institutional diversion (~5pp fewer long nursing-home stays), against a documented partial-offset cost penalty. A6 (PACE) has real autonomy evidence and a real cost penalty. A1 (federal social insurance) is the weakest row with any evidence at all — two below-neutral cells and no positive evidence anywhere, because the record's international instances are all under documented cost and workforce strain inside the measurement window and the domestic instances are one repealed-before-launch design failure and one program with weeks of payout history. A3, A8 and A10 come out identically flat at 3 across all five objectives — I say that plainly rather than manufacture differentiation: A3's 3s are mostly evidenced washes (money did raise wages; it did not move turnover, vacancy or waitlists, and ws03 forbids resolving it), whereas A8's and A10's 3s are pure absence of evidence. Those are different findings sharing a number. The board is mostly flat: 37 of 50 cells sit at exactly 3, and no instrument in this record has evidence bearing on institutional quality in institutional or residential settings except A7 (which washes) and A6 (which washes).

4. Evidence-coverage count

Cells scored on cited evidence (i.e. moved off 3), out of 5:

Architecture Non-3 cells Which 3s that are evidenced-wash (evidence exists, does not resolve)
A1 2 CC, WS 0
A2 3 CC, WS, AA 0
A3 0 2 (WS, AA)
A4 3 CC, CR, AA 0
A5 1 CR 0
A6 2 CC, AA 1 (IQ)
A7 1 WS 1 (IQ)
A8 0 0
A9 1 CC 0
A10 0 0

13 of 50 cells rest on cited evidence of effect or failure; 4 further cells are evidenced washes; the remaining 33 are unevidenced neutrals.

5. What you could not score, and what would change my mind

Structurally unscorable in this record:

What would change my mind, per cell family:

← All Elder care research documents Sources digest Read the whitepaper