GUBMENTPlain talk · policy frontier
Filings / Mental health / Sources / Structurally blinded re-score & reconc
GBMT-11 · Research record · No. 11

Structurally blinded re-score & reconciliation log (GBMT-11)

mental-health/research/ws09-rescore-log.md
This is a working research document from the mental health filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 11, not a summary of it.

Date: 2026-08-11 Method: Per M6/M10 — structural blinding via scripts/batch-rescore.py. Raw scorer output, full packet manifest, and brief: ws09-blind-scores-2026-08-11.md, ws09-blind-brief-2026-08-11.md. Batch: msgbatch_013vUDQ6mt6JV8mz2H95zfsx (model claude-opus-5). Supersedes the pass-1 post-red-team board in ws09-scorecard.md as the filing's independence claim for numeric ranks. Red-team cell moves (#2 O1 5→4, O2 4→3; #6 O1 4→3) are independently re-derived below — the blind scorer matched those three cells outright without seeing the red-team log.

Scale note. Unlike elder care's 2026-08-10 pass, this filing already used the symmetric evidence floor on pass 1 (below-neutral requires cited failure; absence → 3). The blind brief's anchors match that floor. Divergence here is about evidence application, not a mid-pass instrument rewrite.


Headline: CCBHC / crisis leadership holds; mid-board reorders

The post-red-team claim that survived to this pass —

#5 leads W1; #2 CCBHC leads W2 and W3; #10 last under every weighting

survives reconciliation. Seven cells move; three differences are kept at the published value (one evidence keep, two judgment splits). The board does not flip its top pair or its bottom pole.

What does not survive is the impression that mid-pack differentiation was settled. After reconcile, #3 IMD and #7 AOT are numerically identical (3,4,3,3,4), and #8 CoCM and #9 workforce liberalization are both all-neutral. AOT's W2 near-top placement was an O1 stretch; bed-rebuild's O2=4 was problem documentation scored as instrument effect.


Reconciliation

The blind scorer differed from the published board on 10 of 50 cells and matched it on 40. On reconciliation 7 of those 10 are corrected onto the published board; 3 are kept at the published value (1 scale-text keep on #3 O2; 2 judgment splits on #9 O1/O3). Every cell is accounted for. No difference was ≥2 points.

A. Corrections adopted from the blind scorer

Cell Published Blind Resolved Basis
#4 Bed rebuild — O2 4 3 3 Published O2=4 rested on NRI shortage reports, TAC Construct A, and Australia as need documentation (ws03, ws07). The blind correctly applies the house rule: problem size is not instrument effect, and no evaluation of a bed-rebuild / state-hospital reinvestment program appears in the record. Need evidence cannot hold a cell above 3.
#4 Bed rebuild — O5 3 4 4 Published held O5 at unevidenced-neutral. Blind cites NRI Use of State Psychiatric Hospitals, 2025 and TAC staffed-bed censuses: state mental health agencies already operate state psychiatric hospitals at material scale with measured reporting — O5 anchor 4's "closely related instrument already run." Held at 4 not 5 because expansion delivery is unmeasured.
#6 Parity ERISA — O5 2 3 3 Published O5=2 treated DOL OIG 09-25-001 investigator thinness (<1 per ~16,472 plans) as pure capacity failure for census-style exams (ws06). Blind nets that against EBSA's demonstrated CAA comparative-analysis program at scale (corrections since Feb 2021; MA/NY state exam machinery on fully insured). O4 already carries the KC3 / non-census limit. Demonstrated spot-check operation + capacity shortfall = evidenced wash → 3, not 2.
#7 AOT — O1 4 3 3 Published O1=4 stretched Swartz hospital-admission OR 0.77 into "realized access." O1's own anchors score appointment offer, wait-to-evaluation, or crisis connection (ws07). Swartz / Laura's Law measure hospitalization, homelessness, and arrest — engagement outcomes, not O1 proxies. Service intensification (ICM/ACT) is a confound, not an access metric.
#7 AOT — O2 3 4 4 Published held O2 at 3 as "engagement ≠ bed-capacity add." Blind scores the same Swartz stack on O2's throughput language for the selected SMI band the row is specified to score: initial 6-month order → admission OR 0.77; renewal OR 0.59 (ws07). Not 5: pre/post, selection, regression-to-mean, co-delivered ACT. Does not travel to population acute capacity.
#8 CoCM — O5 4 3 3 Published O5=4 credited CoCM billing codes and FQHC BH expansion as operating surfaces. Blind nets that against ws05's own binders: same-day billing bans and team-payment rules that prevent integrated encounters in some states (ws05). No CoCM outcome evaluation is in the packet. Codes-exist ≠ demonstrated scaled delivery of the instrument as specified → wash at 3.
#10 Do-nothing — O1 2 3 3 Published O1=2 scored only the outpatient status-quo failure (Bishop / Brahmbhatt / phantom networks). Blind notes O1 covers appointment or crisis response, and the do-nothing row includes the current 988 funding path where answer-rate and routed-volume gains are measured (ws04). Outpatient-poor + crisis-improving nets to evidenced wash. O2–O4 2s still hold the row last.

B. Kept at published value — scale text, not aesthetics

Cell Published Blind Kept Why
#3 IMD repeal/waiver — O2 4 3 4 Blind reads statute+CMS financing relief against McBain's Construct B null as an evidenced wash. O2's authored band 4 explicitly includes "admit-financing geography" (ws03; scale in scorecard / 00-scales.md). Relieving the documented Medicaid purchase constraint for adult IMD stays is partial evidence on that clause. McBain null is about bed counts — already disclosed in the published basis ("not a proven bed-rebuild") — and does not erase the financing-geography move. Blind's wash is defensible; the scale's own text carries the published 4.

C. Judgment splits kept — published value stands, blind reading recorded

Cell Published Blind Kept Why
#9 Workforce liberalization — O1 3 2 3 Blind cites ws02's licensed≠available stack and the scorecard-implications line ("architectures that only grow licensed headcount… score poorly on O1/O3") as channel failure. Blind's own disclosure: no compact / supervision-ratio / peer-billing instrument is evaluated in the base, and peer-specialist Medicaid billing is a payer-participation lever the headcount critique does not fully reach. Pass-1 red-team Attack 4 already refused inventing below-neutral cells from mediating-channel evidence. Symmetric floor: below-neutral requires cited failure of the instrument.
#9 Workforce liberalization — O3 3 2 3 Same split on the same citations. Bishop / Brahmbhatt diagnose the status-quo wedge (#10's O3=2); they do not evaluate interstate compacts or peer-billing take-up.

D. Unchanged — blind scorer independently matched the published value

Forty cells matched outright, including the three red-team amendments:


Reconciled matrix

# Architecture O1 O2 O3 O4 O5
1 Medicaid behavioral rate floor + admin simplification 3 3 4 3 4
2 CCBHC expansion as default safety-net model 4 3 4 3 4
3 IMD repeal or broad MH IMD waiver 3 4 3 3 4
4 Bed rebuild / state hospital reinvestment 3 3 3 3 4
5 988 + mobile crisis + stabilization continuum 4 4 3 3 4
6 Parity enforcement with ERISA teeth 3 3 3 4 3
7 Assisted outpatient treatment expansion 3 4 3 3 4
8 Primary-care BH / Collaborative Care 3 3 3 3 3
9 Workforce liberalization 3 3 3 3 3
10 Do-nothing comparator 3 2 2 2 4

Bold = moved vs pass-1 post-red-team. #3 and #7 are identical rows. #8 and #9 are identical all-3s.


Rankings under three weightings (recomputed)

W1 — Medicaid-SMI enrollee (0.30 / 0.30 / 0.20 / 0.15 / 0.05)

Rank Architecture Weighted score
1 #5 988 + mobile crisis + stabilization 3.65
2 #2 CCBHC expansion 3.55
3 (tie) #3 IMD repeal/broad waiver · #7 AOT 3.35
5 #1 Rate floor + admin simplification 3.25
6 #6 Parity with ERISA teeth 3.15
7 #4 Bed rebuild / state hospital 3.05
8 (tie) #8 CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.40

W2 — Commercially insured parent (0.35 / 0.05 / 0.15 / 0.35 / 0.10)

Rank Architecture Weighted score
1 #2 CCBHC expansion 3.60
2 #5 988 + mobile crisis + stabilization 3.50
3 #6 Parity with ERISA teeth 3.35
4 #1 Rate floor + admin simplification 3.25
5 (tie) #3 IMD · #7 AOT 3.15
7 #4 Bed rebuild / state hospital 3.10
8 (tie) #8 CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.55

W3 — State BH commissioner (0.10 / 0.20 / 0.30 / 0.10 / 0.30)

Rank Architecture Weighted score
1 #2 CCBHC expansion 3.70
2 (tie) #1 Rate floor + admin · #5 Crisis continuum 3.60
4 (tie) #3 IMD · #7 AOT 3.50
6 #4 Bed rebuild / state hospital 3.30
7 #6 Parity with ERISA teeth 3.10
8 (tie) #8 CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.70

Sensitivity / what moved

Robust across blind + reconcile:

Not robust / withdrawn as worded:

New board facts, stated plainly:


What this pass does not settle

Effects applied

← All Mental health research documents Sources digest Read the whitepaper