Date: 2026-08-10 Method: Per M6/M10 — structural blinding via scripts/batch-rescore.py. Raw scorer output, full packet manifest, and every redaction: ws11-blind-scores-2026-08-10.md. Supersedes the 2026-08-03 pass-2 re-score (red-team-log, "Independent re-score and reconciliation") as this filing's independence claim. That pass is not withdrawn — all three of its corrections are independently re-derived below — but it was an in-workflow agent instructed not to read the scorecard, with filesystem access throughout. It is not the structurally blinded pass M6 requires, and the filing has been claiming otherwise on its public pages.
Why this was owed. On 2026-08-06 the series ran structurally blinded re-scores for housing, media and drugs and stated a cross-filing meta-finding on all three scorecards: most apparent ranking differentiation was evidence-density, not measured instrument difference. Elder care was deliberately excluded from that pass because "elder care's pass-2 re-score already existed" (commit 40b7372). It therefore never received the correction the other three got. This log is that correction.
Headline: the board's clearest negative finding is not robust
The published scorecard's most-repeated claim — stated twice in ws11-scorecard.md, again in the red-team log, and on both public pages — is:
Private LTC insurance reform still ranks last under all three weightings, with no exceptions — the clearest, most stable negative finding in the whole scorecard.
That row scores 2/1/1/1/2. Exactly one of those five cells rests on cited evidence — CC=2, from Milliman's "little potential to generate savings" — and ws04 records that the Milliman PDF was never retrieved, so even that cell comes from secondary reporting of a study nobody on this project has read. The other four cells rest on nothing at all. The record contains no finding, in either direction, on reinsurance and institutional quality, workforce, caregiver relief, or autonomy. They sit at 1 and 2 because the record is silent and the published scale lets silence sit below neutral.
That much is a fact about the record, and it holds under any scoring convention. What follows from it depends on the convention, and this log reports both readings rather than the one that makes the finding bigger:
| Reading | Where private LTC reform ranks |
|---|---|
| Published scale kept, only evidence-based corrections applied | Last under all three weightings — the claim survives |
| Corrected symmetric scale (this pass) | Seventh of eight under all three — the claim falls |
So the honest verdict is not "the claim is false." It is: the claim is not robust. It survives only under a scoring convention in which never having researched an instrument counts as evidence against it. The filing presented it as its most stable finding. It is in fact the finding most sensitive to the scale, because it is the row with the least evidence behind it.
The instrument defect, stated plainly
The published anchored scale is asymmetric, and the asymmetry runs one way:
| Anchor | Published definition | Requires evidence? |
|---|---|---|
| 5 | Direct causal or strong quasi-experimental evidence | Yes |
| 4 | Positive evidence from an analogous or partial natural experiment | Yes |
| 3 | Mixed/indeterminate — plausible mechanism, but real countervailing evidence | Yes (countervailing) |
| 2 | Weak/theoretical case only — no empirical support found | No |
| 1 | No credible evidence of a positive effect; the mechanism plausibly cuts the wrong way | No |
A cell with no evidence in either direction cannot reach 4 or 5, fails 3's "real countervailing evidence" test, and lands at 2 — or at 1, if the scorer can articulate a mechanism argument. Anchors 4 and 5 demand evidence; anchors 1 and 2 accept its absence. An architecture nobody researched is therefore scored down, and a mechanism hunch is sufficient to reach the floor.
This is the same defect the housing re-score named on 2026-08-06 ("the evidence floor was enforced in one direction only"). It is a defect in the instrument, not a disagreement between scorers.
The change made, disclosed rather than buried. The blinded brief widened anchor 3 to cover "no evidence in the record either way", and narrowed anchor 1 to require documented adverse effect rather than a plausible-sounding mechanism. Anchors 2, 4 and 5 were left as authored. This is a material change to the scoring instrument, made by this pass, and it drives most of the movement below. A reader who rejects it should read the sensitivity column above and the "scale-independent corrections" section, which together show exactly how much of this log survives without it.
Reconciliation
The blind scorer's matrix differed from the published board on 26 of the 40 cells across the eight ranked rows, and matched it outright on 14. On reconciliation 25 of those 26 were corrected and 1 was kept at the published value as a logged judgment split — so the published board changes on 25 of 40 cells. Nine of the differences were ≥2 points. Every cell is accounted for.
A. Corrections that hold regardless of the scale change
These do not depend on the widened anchor 3. Anyone who rejects the scale change still owes these.
| Cell | Published | Resolved | Basis |
|---|---|---|---|
| A2 HCBS de-capping — WS | 2 | 4 | ws13 (2026-08-06): a 2026 Health Services Research stacked DID over ACS 2005–21 finds BIP adoption raised the HCBS workforce 13.24%, with no significant institutional-workforce offset — net capacity, not shifted staff. The published 2 predates ws13 and rests on ARPA §9817, whose verdict is INDETERMINATE, not negative. Held at 4 not 5 on ws13's own boundary: this evidences the bundled, target-bound BIP design, not a generic claim that HCBS money creates staff. That boundary is carried into the basis column so the row cannot be read as the generic claim. ws13 stated the row "must not be re-scored without the required independent re-score" — this is that re-score |
| A7 federal AL standards — WS | 3 | 4 | The 2025 Health Affairs 22-state panel (mandates raised direct-care staffing ~5%; CNAs +5.7%, LPNs +7.5%; labour costs rose less than revenue; margins unchanged; no closure effect) is a partial natural experiment in an analogous setting — anchor 4's own definition. The pass-2 re-score moved this cell 1→3 for exactly this evidence and stopped one step short. Scoped: SNF not AL, mandates up to ~4.1 HPRD |
| A7 federal AL standards — IQ | 2 | 3 | Real evidence both ways that nets out — which is the published anchor 3, not the widened one. For: the Health Affairs +5% staffing result; GAO-18-179 and GAO-26-107884 documenting a persistent, government-verified incident-reporting gap (only 22 of 48 states could report critical incidents; only 34 of 48 published any), with a CMS reporting fix finalised May 2024 but not operative until 2027. Against: the nearest federal analogue was killed twice over (vacated N.D. Tex. and N.D. Iowa, appeals withdrawn, ~10-year congressional moratorium, rescinded Feb 2 2026) leaving essentially no compliance record; and NY's pass-through-style law shows ~400 violating facilities with essentially no penalties for years |
| A1 federal LTC insurance — CR | 4 | 3 | Anchor 4 requires an analogous natural experiment. The published basis column itself concedes the row is "plausibility, not outcome evidence yet": WA Cares has ~zero months of payout data and no caregiver-burden outcome exists for Germany, Japan or the Netherlands anywhere in the record. Under the published scale this cell would fall to 2, not 3 — the corrected scale is the more generous of the two here |
| A1 federal LTC insurance — AA | 4 | 3 | Same. No aging-in-place outcome for any instance of this instrument class appears in the record. Same note: the published scale would put this at 2 |
These two above-neutral cells are the only unevidenced cells the published board scored up. It was markedly more disciplined upward than downward — which is the shape of the defect.
B. Corrections that depend on the widened anchor 3
Below-neutral cells with no cited evidence of failure in either direction, returned to neutral.
| Cell | Published | Resolved | Basis |
|---|---|---|---|
| A1 federal LTC insurance — IQ | 2 | 3 | Nothing links a payroll-tax LTC benefit to institutional quality in any of the five named instances |
| A2 HCBS de-capping — IQ | 2 | 3 | Diversion from nursing homes is not evidence about quality inside them |
| A3 wage floor — CC | 2 | 3 | No source measures payer or system cost under a wage floor; Health Affairs' margin result is facility finance, not payer cost |
| A3 wage floor — CR | 2 | 3 | No evidence on family-caregiver outcomes from wage pass-through |
| A4 Cash & Counseling — IQ | 2 | 3 | A community instrument, never evaluated in institutional or residential settings |
| A4 Cash & Counseling — WS | 2 | 3 | ws08 calls the agency-workforce interaction untested by the researchers who would know; ws13 instructs against inferring it. Honoured, not resolved |
| A7 federal AL standards — CC | 2 | 3 | No evidence on payer cost from setting standards |
| A7 federal AL standards — CR | 1 | 3 | No evidence on caregiver outcomes from AL or SNF standards. The 1 was mechanism reasoning |
| A7 federal AL standards — AA | 2 | 3 | Residential standards never evaluated against community tenure or resident autonomy |
| A8 OAA/NFCSP — CC | 2 | 3 | The pass-2 correction (4→2) rightly rejected treating programme scale as cost containment, but landed one step past neutral: no file evaluates NFCSP's cost effect in either direction |
| A8 OAA/NFCSP — IQ | 1 | 3 | No evidence |
| A8 OAA/NFCSP — WS | 1 | 3 | No evidence |
| A8 OAA/NFCSP — AA | 2 | 3 | Home-delivered meals supporting community tenure is a mechanism argument only |
| A9 private LTC reform — IQ | 1 | 3 | No evidence |
| A9 private LTC reform — WS | 1 | 3 | No evidence |
| A9 private LTC reform — CR | 1 | 3 | ws04's own red-team correction (#8) rescoped its finding to savings, which is a CC finding. No caregiver-relief evidence exists for this design |
| A9 private LTC reform — AA | 2 | 3 | No evidence on where privately insured recipients receive care |
| A10 cash comparator — CC | 2 | 3 | No evaluation of unconditional cash to older adults or families anywhere in the record |
| A10 cash comparator — IQ | 1 | 3 | No evidence |
| A10 cash comparator — WS | 1 | 3 | No evidence |
Twenty cells. Not one was corrected because new evidence arrived. They were scored below neutral on mechanism reasoning and absence of research.
C. Reasoning replaced, score unchanged
| Cell | Value | What changed |
|---|---|---|
| A3 wage floor — WS | 3 | The published basis justifies this partly via the Health Affairs 22-state panel. That study evaluates staffing mandates, not wage pass-through; it belongs to A7 and has been moved there. A3's WS=3 now stands as an evidenced wash on its own evidence: ARPA §9817 moved wages substantially (Colorado 12.41→ 18; national median 13.07→16.77) while every state in KFF's 2023 survey still reported shortages and Pennsylvania turnover held at 44–65%. Same number, now the number its own evidence supports |
| A1 federal LTC insurance — CC | 2 | Published basis was plausibility. Now evidenced: German contribution rates rose again for 2025 with federal loans needed to hold 2026 steady; the Dutch Wlz's own cost-containment goal is partly undermined by perverse incentives; Japan is proposing copays toward 30%; HHS's 19-month CLASS review could not price the design into solvency under any fix tried |
| A1 federal LTC insurance — WS | 2 | Now evidenced: German direct-care vacancies rose 110% in a decade against 45% demand growth and 13% workforce growth; Japan projects a 570,000-worker shortfall by 2040; the Netherlands 266,000 by 2035 |
| A9 private LTC reform — CC | 2 | Unchanged, but the record now states what carries it: one secondary-sourced, never-directly-read Milliman study, scoped to Washington's design. This is the only evidenced cell in the row |
| A4 Cash & Counseling — CC | 1 | Independently re-derived. Both passes record the same tension: AR's cost premium converged to non-significant by year 2 while FL/NJ persisted. The blind scorer called this "my most aggressive score". Kept at 1 — the adverse finding is RCT-grade and persisted in two of three states — with the year-2 convergence now stated in the basis rather than omitted |
D. Judgment split kept — published value stands, blind reading recorded
| Cell | Published | Blind | Kept | Why |
|---|---|---|---|---|
| A4 Cash & Counseling — AA | 4 | 5 | 4 | The objective bundles autonomy with aging-in-place. The RCT supplies experimental evidence on the autonomy half (satisfaction 60.8% vs 42.7% in AR; self-direction closed a gap where controls never received authorised services) but measures neither institutionalisation nor community tenure. The blind scorer flagged 4 as defensible on exactly this ground. Half an objective evidenced at 5 does not make the objective a 5 |
E. Unchanged — blind scorer independently matched the published value
A1 CC (2), A1 WS (2), A2 CC (2), A2 CR (3), A2 AA (4), A3 IQ (3), A3 WS (3), A3 AA (3), A4 CC (1), A4 CR (5), A8 CR (3), A9 CC (2), A10 CR (3), A10 AA (3).
Fourteen cells matched outright. A8 CR is worth naming: the blind scorer recorded that NFCSP's ~$209M against a 234B–1.01T caregiver-value range "tempts a 2", and held at exactly 3 because programme scale is not evidence of an instrument's effect. That is the pass-2 reconciliation's own reasoning, reached independently.
The one cell the board cannot score, and it is the record's best evidence
The strongest quality-and-safety finding in the entire elder-care record is the Gupta/Howell/Yannelis/Gupta within-facility IV estimate: +11% mortality and ~50% higher antipsychotic use in private-equity-owned nursing homes, across
7M Medicare patients. It attaches to no row on this board.
ws06-09explicitly recommends PE-ownership disclosure and PE-specific oversight "as a candidate architecture independent of the staffing-ratio question" — and it was never added to the candidate list. The blind scorer surfaced this independently: "If PE oversight were row A11, it would be the only row with a 5 on IQ."
A scope finding, not a scoring finding, and not fixed here — adding an architecture needs pre-registered criteria and its own pass. Recorded because the board's IQ column is nine-tenths empty while the record's best IQ evidence sits outside it.
Reconciled matrix
| Architecture | CC | IQ | WS | CR | AA | Sum |
|---|---|---|---|---|---|---|
| Federal LTC social insurance (WA Cares nationalized) | 2 | 3 | 2 | 3 | 3 | 13 |
| Medicaid HCBS waiver expansion / de-capping | 2 | 3 | 4 | 3 | 4 | 16 |
| Direct-care wage floor tied to funding | 3 | 3 | 3 | 3 | 3 | 15 |
| Cash & Counseling / self-direction expansion | 1 | 3 | 3 | 5 | 4 | 16 |
| Credit for Caring Act | — | — | — | 2–3 | — | not ranked |
| PACE expansion | — | — | — | — | — | not ranked (see below) |
| Federal minimum standards for assisted living | 3 | 3 | 4 | 3 | 3 | 16 |
| OAA/NFCSP expansion | 3 | 3 | 3 | 3 | 3 | 15 |
| Private LTC insurance market reform (this reinsurance design) | 2 | 3 | 3 | 3 | 3 | 14 |
| Cash as the comparator (per M7) | 3 | 3 | 3 | 3 | 3 | 15 |
30 of the 40 ranked cells now sit at 3. That is the honest state of this evidence base, reported as flatness rather than dressed up as differentiation.
Rankings under three weightings
Simple sum of the five objectives, then the named objective weighted ×3. Credit for Caring and PACE remain excluded. Recomputed in full.
Workforce-first (WS×3): HCBS de-capping (24) = Federal AL standards (24) > Cash & Counseling (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > Private LTC insurance reform (20) > Federal LTC insurance (17)
Caregiver-relief-first (CR×3): Cash & Counseling (26) > HCBS de-capping (22) = Federal AL standards (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > Private LTC insurance reform (20) > Federal LTC insurance (19)
Cost-containment-first (CC×3): Federal AL standards (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > HCBS de-capping (20) > Cash & Counseling (18) = Private LTC insurance reform (18) > Federal LTC insurance (17)
Sensitivity — the same board under the published scale (per protocol S5)
Applying only the section-A corrections and leaving the published asymmetric scale intact:
WS×3: HCBS de-capping (23) > AL standards (20) > Wage floor (19) > Cash & Counseling (18) > Federal LTC insurance (16) > Cash (12) > OAA (11) > Private LTC reform (9) CR×3: Cash & Counseling (24) > HCBS de-capping (21) > Federal LTC insurance (18) > Wage floor (17) > Cash (16) > OAA (15) > AL standards (14) > Private LTC reform (9) CC×3: HCBS de-capping (19) > Wage floor (17) > Federal LTC insurance (16) = Cash & Counseling (16) = AL standards (16) > Cash (14) > OAA (13) > Private LTC reform (11)
What is robust across both readings:
- The wage floor no longer leads workforce weighting. Published pass 2 had it first at 19; it is fourth under the corrected scale and third under the published one. HCBS de-capping leads workforce weighting under both readings. This is the most consequential scale-independent change on the board.
- HCBS de-capping rises sharply under both readings, on ws13's BIP evidence.
- Federal AL standards rises under both readings.
- Cash & Counseling does not win under every weighting — the pass-2 correction survives both readings.
- Cash & Counseling leads caregiver-relief weighting decisively under both.
What is contingent on the scale change: the withdrawal of "private LTC reform ranks last with no exceptions"; federal LTC social insurance ranking last; the three-way tie at 21; and the characterisation of the board as flat.
What changed, what survived
Not robust — "private LTC insurance reform ranks last under all three weightings, with no exceptions." See the headline and sensitivity sections. The claim survives the published scale and falls under the corrected one; what holds under both is that four of its five cells rest on no evidence at all. It cannot continue to be published as "the clearest, most stable negative finding in the whole scorecard" — it is the least stable one.
Withdrawn — "it ties or trails nationalized WA Cares insurance and the wage floor" (of Cash & Counseling). Under the corrected board it beats nationalized WA Cares insurance under all three weightings and beats the wage floor under two of three. What survives is the narrower and more important half: it does not win under every weighting — third under workforce, joint-sixth under cost, held down by its own RCT-documented cost increase.
Survives and strengthens — Cash & Counseling leads decisively wherever caregiver relief is weighted heavily. Its margin widens from 24-vs-22 to 26-vs-22. It remains the only architecture anchored by an actual randomized trial, and the only one carrying RCT-grade evidence against itself on another objective. Still the filing's best single finding.
Survives — all three pass-2 corrections, each independently re-derived from the evidence alone by a scorer who could not see them. The 2026-08-03 pass was under-blinded, not wrong.
New — federal minimum standards for assisted living is the largest single move on the board, from last-or-near-last in pass 1 and mid-table in pass 2 to first under cost weighting and joint-first under workforce weighting. Most of that is the removal of unevidenced 1s and 2s; the rest is the Health Affairs staffing-mandate panel finally being applied to the row it actually evidences. Read this as the board admitting it knows very little about eight of these instruments, not as a discovery that AL standards are excellent.
New — the filing's sequencing prior gets no support from its own board. §12 orders workforce first; the wage floor is joint-fourth and tied with the do-nothing cash comparator. The sequencing argument rests on the workstream reasoning, not on the scorecard, and the record should say so rather than let the two appear to corroborate each other.
New — a three-way tie at 21 under every weighting between the direct-care wage floor, OAA/NFCSP expansion, and unstructured cash. Three instruments the filing treats as very different are indistinguishable on this evidence base. Stated plainly rather than differentiated into a false ordering.
New — federal LTC social insurance ranks last under all three weightings on the corrected board. Unlike the row it displaced, this one is evidenced: real adverse findings on cost and workforce for every international instance, plus a domestic design-stage failure. It ranks last on evidence, which is what the published bottom rank claimed to be and was not.
PACE — scored blind for the first time, and still not entered
PACE has never been scored. ws12 carries a provisional row it explicitly declines to enter. The blind scorer, working from an evidence-only extract of ws12 with that provisional row and its sequencing recommendation withheld, produced an independent read:
| Cell | ws12 provisional | Blind (independent) | Agreement |
|---|---|---|---|
| CC | 2 | 2 | corroborated |
| IQ | 4 | 3 | contested |
| WS | 3 | 3 | corroborated |
| CR | 3 | 3 | corroborated |
| AA | 4 | 4 | corroborated |
Four of five corroborated from the evidence alone. IQ is contested downward: the blind scorer read CMS/Abt's favourable six-month result against its own narrowing at later follow-up, its acknowledged selection bias, and the fact that PACE is a community instrument that does not speak to quality inside institutional settings — and called it an evidenced wash at 3.
PACE is still not entered on the published board. Deviations #22, #24, #26 and #28 govern: the M3 integration criteria were written after the evidence pass, and the prospective check returned indeterminate on all four. A re-score does not satisfy those criteria. This log records the independent read so a future integration pass starts from two derivations rather than one.
What this pass does not settle
- Two re-scores is "twice checked," not final. M10 is explicit and GBMT-10 is the precedent: two independent blind re-scores of the crypto board moved disjoint cell sets. A third pass would very likely move cells these two agreed on.
- The scale change is doing most of the work, and the sensitivity section shows exactly how much. A reader who rejects it should treat the published board as intact apart from section A, and read the rest as a disagreement about instruments rather than about evidence.
- The board's flatness is now the finding. With 30 of 40 cells at neutral, the weightings separate very little and several "rankings" are ties. Reporting rank order at all risks implying more resolution than exists.
- Credit for Caring has a full blind row (3/3/3/2/3) that is not entered: the pass-2 decision to split it out and leave it unranked was a red-team judgment this pass has no mandate to reverse.
- The record's best quality evidence still has no row. See the PE-ownership note above.
- This is a re-score, not a fact-check. It takes the evidence base as given. Whether the underlying findings are true is Phase 1 of the Verification Protocol, which has not yet run on this filing.
Effects applied
- ws11-scorecard.md → reconciled matrix, recomputed rankings, rewritten rank-stability section; prior values in git history.
- deviations-log.md → entry #30.
site/elder-care/index.html→ scorecard table, Part 8 prose, sidebar stamp, honesty box, receipts line.site/elder-care/sources/index.html→ §11 digest, red-team re-score paragraph, deviations digest, receipts line.HANDOFF.md→ elder-care blind-re-score open item discharged; series table note ³ rewritten.