Per M6: anchored ordinal scale authored first, every cell cites a workstream finding, ranked under multiple objective weightings with rank-stability reported.
Three scoring passes, not one. Pass 1 was a single scorer. Pass 2 added an independent re-score and reconciliation (3 cells moved ≥2 points) plus red-team corrections — but that scorer was instructed not to read this file while having filesystem access throughout, which is instruction-level blinding, not the structural blinding M6 requires. Pass 3 (2026-08-10) closed that gap: a scorer with no tools at all, working from an evidence base inlined into a single request, which could not reach this file by any route. 25 of the 40 ranked cells moved (the blind scorer differed on 26; one was kept at its published value as a logged judgment split). Full reconciliation, including a sensitivity analysis showing which findings survive under the old scale and which do not: ws11-rescore-log.md. Pass 2's resolution log remains at red-team-log.md; all three of its corrections were independently re-derived in pass 3.
The scale was changed in pass 3, and that change drives most of the movement. See "Anchored scale" below and the rescore log's sensitivity section. Read both before citing any ranking here.
Anchored scale (1–5) — amended in pass 3, with the original preserved
As authored (passes 1–2):
1 = No credible evidence of a positive effect in this domain; the mechanism plausibly cuts the wrong way. 2 = Weak/theoretical case only — no empirical support found, or a specific finding undermines it. 3 = Mixed/indeterminate — plausible mechanism, but real countervailing evidence from this session's research. 4 = Positive evidence from an analogous or partial natural experiment/precedent — correlational, not causal. 5 = Direct causal or strong quasi-experimental evidence (e.g., an RCT) supporting a real effect.
As amended (pass 3, 2026-08-10) — anchors 2, 4 and 5 unchanged:
1 = The evidence base contains direct evidence of an adverse effect on this objective. 3 = Mixed/indeterminate — plausible mechanism with real countervailing evidence, or no evidence in the record either way.
Why. The original scale is asymmetric: anchors 4 and 5 both require cited evidence, while anchors 1 and 2 accept its absence (anchor 2 is "no empirical support found"; anchor 1 needs only a mechanism that "plausibly cuts the wrong way"). Under it, an architecture nobody researched is scored down and can never be scored up — so the bottom of this board measured research effort rather than instrument quality. That is the same defect the housing re-score named on 2026-08-06 ("the evidence floor was enforced in one direction only"), and the same correction applied to housing, media and drugs. Elder care was excluded from that pass because a re-score already existed; this is the correction it missed.
This amendment is load-bearing and is disclosed as such. Roughly two thirds of pass 3's corrected cells are below-neutral cells with no cited evidence of failure returning to neutral. The rescore log publishes the full sensitivity: which findings hold under both scales, and which exist only under the amended one.
Scores (pass 3 — reconciled)
Objectives (per §1.5): CC cost containment, IQ institutional quality & safety, WS direct-care workforce sustainability, CR family-caregiver relief, AA aging-in-place/autonomy.
| Architecture | CC | IQ | WS | CR | AA | Basis |
|---|---|---|---|---|---|---|
| Federal LTC social insurance (WA Cares nationalized) | 2 | 3 | 2 | 3 | 3 | CC and WS are the row's only evidenced cells, and both are adverse: German contribution rates rose again for 2025 with federal loans needed to hold 2026 steady, the Dutch Wlz's own cost-containment goal is partly undermined by perverse incentives, Japan is proposing copays toward 30%, and HHS's 19-month CLASS review could not price the design into solvency under any fix tried; on workforce, Japan projects a 570,000-worker shortfall by 2040 and the Netherlands 266,000 by 2035 (ws10, phase0 §4). Basis narrowed on Phase 1 verification (2026-08-10, verification-log): this cell previously also cited "German vacancies rose 110% in a decade against 13% workforce growth" to NBER w31870, which contains no such series — the figures are withdrawn pending a citable source, and the German 2026 rate-loan claim behind the CC cell rests on a consumer-finance website. The cell value does not move (2 either way), but this row's claim to rank last on evidence rather than on absence of it is weaker than the re-score log states, and should not be repeated until the German sourcing is fixed. CR and AA corrected 4→3 in pass 3 — anchor 4 requires an analogous natural experiment and none exists: WA Cares has ~zero months of payout data and no caregiver-burden or aging-in-place outcome for Germany, Japan or the Netherlands appears in the record. Mandatory participation still avoids CLASS's adverse-selection death, but that is a design argument, not an outcome |
| Medicaid HCBS waiver expansion/de-capping | 2 | 3 | 4 | 3 | 4 | WS corrected 2→4 in pass 3. A 2026 Health Services Research stacked DID (ACS 2005–21) finds Balancing Incentive Program adoption raised the HCBS workforce 13.24% with no significant institutional-workforce offset — net capacity, not shifted staff (ws13). The prior 2 rested on ARPA §9817, whose verdict is INDETERMINATE, not negative. Scope, which the score does not license waiving: this evidences the bundled, target-bound BIP design, not a generic claim that HCBS money creates staff — a 2025 state-year study found no significant wage association with HCBS spending. Scope tightened on Phase 2 verification (2026-08-10, steelman-log): the same paper evaluates two programs and finds no significant effect for Community First Choice (§1915(k)) — 1.51%, 95% CI −12.77% to 15.79% — which this column previously omitted while citing the paper's BIP arm. CFC is the closest enacted analogue to this row and to the Cash & Counseling row, and it carries a 6-percentage-point FMAP bonus, so the null is directly on point. The cell holds at 4 on BIP, but this row's evidence now explicitly includes a null result for the statutory version of the same idea. AA=4 on a 2025 BMJ Open generalized DID (~5pp fewer long nursing-home stays for beneficiaries living alone), held below 5 on its own named limits. CC=2: each added HCBS dollar is associated with $0.74 of added total LTSS spending — partial offset only, self-financing unsupported |
| Direct-care wage floor tied to funding | 3 | 3 | 3 | 3 | 3 | CC corrected 2→3, CR corrected 2→3 in pass 3 — no source measures payer or system cost under a wage floor, and none measures caregiver outcomes from one; both were scored below neutral on mechanism. WS=3 is an evidenced wash and its basis is corrected: pass 2 justified it partly via the Health Affairs 22-state panel, but that study evaluates staffing mandates, not wage pass-through, and has been moved to the AL-standards row. On its own evidence, ARPA §9817 moved wages substantially (Colorado 12.41→ 18; national median 13.07→16.77) while every state in KFF's 2023 survey still reported shortages and Pennsylvania turnover held at 44–65% — INDETERMINATE, per ws03 and red-team-log #5 |
| Cash & Counseling / self-direction expansion | 1 | 3 | 3 | 5 | 4 | The only architecture in this table with actual RCT evidence — measurably better satisfaction, less caregiver strain, better self-rated health across all three states (ws08). CC=1 because its own RCT cost data (AR +17%, FL +14%, both significant, year 1) is direct evidence of a cost increase; the premium converged to non-significant in Arkansas by year 2 but persisted in FL/NJ, and the mechanism was closing unmet need rather than inefficiency — it did more of the job, not the same job cheaper. Corrected 3→1 in pass 2 and independently re-derived in pass 3. IQ and WS corrected 2→3 in pass 3: never evaluated in institutional settings, and both ws08 and ws13 explicitly instruct that the agency-workforce interaction is untested and must not be inferred. AA held at 4 as a logged judgment split — the RCT evidences autonomy at 5 but measures neither institutionalisation nor community tenure |
| Credit for Caring Act | — | — | — | 2–3 | — | Split from the row above on red-team review (#17) — different eligible population, no shared administrative machinery, no empirical test of the bill itself (no CBO/JCT score exists; stalled in committee every Congress since 2016). Pass 3's blind scorer produced a full row (3/3/3/2/3) and disclosed that its CR=2 scores a decade of non-advancement as delivery failure rather than any finding about the design's effect. Not entered — reversing pass 2's split-and-unrank decision is outside a re-score's mandate. Recorded in ws11-rescore-log.md |
| PACE expansion | — | — | — | — | — | Still not scored on this board, but no longer resting on an unverified seed characterization. ws12 carries a provisional row it declines to enter; pass 3's blind scorer, working from an evidence-only extract with that row withheld, independently corroborated 4 of its 5 cells (CC 2, WS 3, CR 3, AA 4) and contested IQ downward (4→3, an evidenced wash: CMS/Abt's favourable six-month result narrows at later follow-up, has acknowledged selection bias, and does not speak to quality inside institutional settings). Entry still barred by deviations #22/#24/#26/#28 — the M3 integration criteria returned indeterminate on all four, and a re-score does not satisfy them |
| Federal minimum standards for assisted living | 3 | 3 | 4 | 3 | 3 | The largest single move on the board: every cell changed in pass 3. WS 3→4 on the Health Affairs 22-state panel (March 3, 2026 issue; this file previously dated it 2025 — mandates raised direct-care staffing ~5%; CNAs +5.7%, LPNs +7.5%; labour costs rose less than revenue; margins unchanged; no closure effect; eleven of the twenty-two states had no mandate at all across the window) — a partial natural experiment in an analogous setting, which is anchor 4's own definition; pass 2 moved this cell 1→3 for exactly this evidence and stopped one step short. Scoped to SNF, not AL, and to mandates up to ~4.1 HPRD. IQ 2→3 as an evidenced wash under the original scale too: GAO-18-179 documents a government-verified incident-reporting gap (only 22 of 48 states could report critical incidents; only 34 of 48 published any) and GAO-26-107884 documents a separate, still-open federal data gap on what the setting costs ("at least $12 billion in 2024 … likely an undercount") — basis corrected on Phase 1 verification (2026-08-10): this column previously read the two reports as documenting the same incident-reporting gap; the 2026 report is Assisted Living Facilities: Information on Federal Spending and Medicaid Coverage and makes no incident finding — with a CMS incident-reporting fix finalised May 2024 but not operative until 2027, against a federal analogue killed twice over (vacated in two courts, congressional moratorium to September 30, 2034, rescinded Feb 2 2026) and NY's enforcement record (~400 violating facilities, essentially no penalties for years). CC, CR and AA corrected to 3 as unevidenced. Read the rise as the board admitting how little it knows, not as a discovery that this instrument is excellent. The load-bearing risk on this row remains political precedent, not a demonstrated legal one (#22) |
| OAA/NFCSP expansion | 3 | 3 | 3 | 3 | 3 | Four cells corrected in pass 3; the row is now entirely unevidenced. Pass 2 rightly rejected treating the ~209M − vs−116–162B comparison (anchor 12) as a cost-containment finding and moved CC 4→2 — but landed one step past neutral: no file evaluates NFCSP's cost effect in either direction. IQ 1→3, WS 1→3 and AA 2→3 likewise had no evidence behind them. CR held at 3 by both scorers independently, each noting that the tiny funding ratio tempts a 2 and that programme scale is not evidence of an instrument's effect |
| Private LTC insurance market reform (this reinsurance design) | 2 | 3 | 3 | 3 | 3 | Four of five cells corrected in pass 3, and this row carries the filing's most consequential correction. CC=2 stands and is the row's only evidenced cell: Washington's 2015 Milliman feasibility study found this specific state-subsidized reinsurance design had "little potential to generate savings" — suggestive negative evidence from one unread, secondhand-reported state study (ws04, #8). IQ, WS, CR and AA rested on no record evidence in either direction; ws04's own red-team correction rescoped its finding to savings, which is a CC matter. Scoped to this design; a federal reinsurance mechanism with different risk-corridor/subsidy/mandate mechanics is untested |
| Cash as the comparator (per M7) | 3 | 3 | 3 | 3 | 3 | Three cells corrected in pass 3. The record contains no evaluation of unconditional cash to older adults or families at all, so CC 2→3, IQ 1→3 and WS 1→3. Pass 3's scorer explicitly declined to import the Cash & Counseling RCT here, on the ground that the trial cannot decompose the cash component from the counselling and fiscal-intermediary components — and disclosed that declining is as much a judgment call as importing would be. The comparator cannot currently do the job M7 assigns it: five neutral cells beat nothing and lose to nothing |
Rankings under three weightings
Simple sum of unweighted objectives, then the named objective weighted ×3. PACE and Credit for Caring remain excluded. Recomputed in full against the pass-3 reconciled table above.
Workforce-first (WS×3): HCBS de-capping (24) = AL standards (24) > Cash & Counseling (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > Private LTC insurance reform (20) > Federal LTC insurance (17) Caregiver-relief-first (CR×3): Cash & Counseling (26) > HCBS de-capping (22) = AL standards (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > Private LTC insurance reform (20) > Federal LTC insurance (19) Cost-containment-first (CC×3): AL standards (22) > Wage floor (21) = OAA/NFCSP (21) = Cash (21) > HCBS de-capping (20) > Cash & Counseling (18) = Private LTC insurance reform (18) > Federal LTC insurance (17)
Sensitivity (protocol S5). Because pass 3 amended the scale, the same board is also reported under the original scale with only the scale-independent corrections applied. Under that reading: private LTC insurance reform still ranks last under all three weightings; HCBS de-capping leads workforce and cost weighting; Cash & Counseling still leads caregiver relief; the wage floor still does not lead anything. Full figures in ws11-rescore-log.md. Do not cite a ranking from this file without also citing which scale it came from.
Rank-stability finding (rewritten in pass 3)
"Private LTC insurance market reform ranks last under all three weightings, with no exceptions — the clearest, most stable negative finding in the whole scorecard" is withdrawn as stated. It was the opposite of stable. That row scored 2/1/1/1/2, and exactly one of those five cells rested on cited evidence — the CC=2 from a Milliman study that this project never read, only heard reported. The other four were the scale converting absence of research into evidence of failure. Under the amended scale the row places seventh of eight under every weighting; under the original scale it still places last. A finding whose entire rank depends on which scoring convention you pick is not the most stable finding on a board — it is the least. The fact underneath it, true either way and now stated instead: nobody has researched what a reinsurance design does to care quality, workforce, caregiver burden or autonomy, and this scorecard should not have implied otherwise.
"Cash & Counseling ties or trails nationalized WA Cares insurance and the wage floor" is withdrawn. On the reconciled board it beats nationalized WA Cares insurance under all three weightings and beats the wage floor under two of three. What survives — and it is the half that matters — is that it does not win under every weighting: third under workforce, joint-sixth under cost, held down by its own RCT-documented cost increase. Pass 1's "wins regardless of weighting" claim stays withdrawn, for the reason pass 2 gave.
What strengthens: Cash & Counseling leads caregiver-relief weighting decisively, and by a wider margin than before (26 vs 22, up from 24 vs 22). It remains the only architecture on this board anchored by an actual randomized trial, and the only one carrying RCT-grade evidence against itself on a different objective. That is still the most useful thing this scorecard says.
What changes and holds under both scales: the direct-care wage floor no longer leads workforce weighting. Pass 2 put it first at 19. It is now joint-fourth — tied with the do-nothing cash comparator and with OAA/NFCSP — and third under the original scale. HCBS de-capping leads workforce weighting under both readings, on ws13's BIP evidence. §12's sequencing prior (workforce first) therefore draws no support from this board. It may still be right on the workstream argument; it is not corroborated here, and the two should not be read as agreeing.
Federal minimum standards for assisted living moved from last-or-near-last (pass 1) to mid-table (pass 2) to first under cost weighting and joint-first under workforce weighting (pass 3). Three passes, three positions. Most of the pass-3 move is the removal of unevidenced 1s and 2s rather than any discovery about the instrument.
Federal LTC social insurance now ranks last under all three weightings on the amended board — a near-inversion of the published finding. Unlike the row it displaced, this one ranks last on evidence: real adverse cost and workforce findings for every international instance of the class, plus a domestic design-stage failure.
Qualified on Phase 2 verification (2026-08-10, steelman-log S5): this is a one-cell finding, and it ties rather than leads. The WS=2 cell's basis is Japan's projected 570,000-worker shortfall by 2040 and the Netherlands' 266,000 by 2035 — projections of national workforce shortfalls in countries that have LTC social insurance, not estimates of the instrument's effect on workforce. The comparator without the instrument, the United States, has the same shortfall: anchor 2 records PHI's 9.7 million total direct-care openings, 2024–2034. Scoring a cell below neutral on a phenomenon the counterfactual also exhibits is the defect the pass-3 scale amendment exists to prevent. Moving that single cell 2→3 (row 2·3·3·3·3, sum 14) takes the row to 20 under workforce weighting (tied with private LTC reform), 20 under caregiver relief (tied with private LTC reform), and 18 under cost (three-way tie with Cash & Counseling and private LTC reform) — last place shared under all three, never sole. Published as a sensitivity, not entered as the cell's value, because deciding a cell is the filing's call and not a verification pass's. Read alongside the private-LTC withdrawal above: a replacement headline resting on one below-neutral cell is the same shape of claim as the one it replaced.
Separately checked and not disturbed: removing the German figures Phase 1 withdrew (the miscited vacancy series, the consumer-finance-sourced 2026 rate loan) leaves CC=2 supported by the CLASS review, the Dutch Wlz and Japan's copay proposals, and WS=2 supported by Japan and the Netherlands. Neither cell moves and the rank is unchanged on that ground.
A three-way tie at 21 under every weighting between the wage floor, OAA/NFCSP expansion, and unstructured cash. Three instruments this filing treats as very different are indistinguishable on this evidence base. Reported as a tie rather than broken into a false ordering.
What this scorecard cannot yet claim
- The board is mostly flat, and that is now the headline finding about it. 30 of the 40 ranked cells sit at exactly 3. The three weightings separate very little, several "rankings" are ties, and reporting rank order at all risks implying more resolution than this evidence base has.
- Twice re-scored is not final. Per M10, and per GBMT-10's precedent — two independent blind re-scores of the crypto board moved disjoint cell sets. A third pass would likely move cells these two agreed on.
- The scale amendment is load-bearing. Most of pass 3's movement follows from it mechanically. The sensitivity analysis above and in the rescore log is not a footnote; it is how this board should be read.
- PACE and the Credit for Caring Act are still not ranked — both now have independent blind scores recorded in the rescore log, neither is entered, and the reasons are procedural (deviations #22/#24/#26/#28 for PACE; pass 2's split decision for the credit), not evidentiary.
- The record's best quality-and-safety evidence has no row on this board. The Gupta/Howell/Yannelis/Gupta within-facility IV estimate (+11% mortality for compliers, in a 4.2-million-patient short-stay Medicare sample) attaches to PE-ownership disclosure/oversight, which ws06-09 explicitly recommends as a candidate architecture and which was never added to the candidate list. The IQ column is nine-tenths empty while the strongest IQ evidence in the filing sits outside the board. (Figures corrected 2026-08-10 on Phase 2 verification: this bullet was missed by the Phase 1 pass that corrected the same claim everywhere else, and still read ">7M Medicare patients" and "~50% higher antipsychotic use" — the latter appears nowhere in the paper at any vintage. See verification-log finding 1.)
- No row on this board is financed by Medicare, and no dimension can register that an architecture has already been enacted. Both gaps were named on Phase 2 verification (steelman-log, S5 sensitivity 3): CMS's GUIDE model pays Medicare dollars for dementia care management and up to $2,500/year of caregiver respite nationwide from 2024-07-01, and Pub. L. 119-21 §71121 partially enacts the HCBS de-capping row from 2028-07-01 with $150M of implementation funding. Neither can move a cell on a board whose five objectives are cost, quality, workforce, caregiver relief and autonomy. Recorded as recommendations for a future pass; adding rows or dimensions is outside a verification pass's mandate.
- Scores for HCBS de-capping, the wage floor, and federal LTC insurance all still rest partly on the ARPA §9817 finding, which is INDETERMINATE — the scorecard inherits that uncertainty and should not be read as more confident than its inputs.
- This is a re-score, not a fact-check. It takes the evidence base as given. Whether the underlying findings are true is Phase 1 of the Verification Protocol, which has not yet run on this filing.