GUBMENTPlain talk · policy frontier
Filings / Mental health / Sources / §9 Candidate Architecture Scorecard (G
GBMT-11 · Research record · No. 11

§9 Candidate Architecture Scorecard (GBMT-11 Mental Health)

mental-health/research/ws09-scorecard.md
This is a working research document from the mental health filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 11, not a summary of it.

Date: 2026-08-11. Pass 2 — structurally blinded re-score reconciled (ws09-rescore-log.md; raw board ws09-blind-scores-2026-08-11.md). Batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx (claude-opus-5 via scripts/batch-rescore.py). Pass-1 post-red-team values remain in git history and in the rescore log. Method: Per M6 / M10 — anchored ordinal scales authored first (outcome/measured-property anchors only — never instrument-category names); every cell cites ws01ws08 or phase 0; ranked under three explicit weight vectors that sum to 1.0; rank-stability reported. Symmetric evidence floor: below-neutral requires cited evidence of failure or harm; absence of evidence → 3, not 1 or 2 (elder-care / housing scale lesson). Blind differed on 10/50 cells; 7 corrected, 3 kept (1 scale-text keep; 2 judgment splits).

Kill-condition posture baked into scoring (not footnotes):

Red-team cell amendments (2026-08-11, independently matched by blind): #2 O1 5→4, O2 4→3; #6 O1 4→3. Full attacks in ws09-red-team-log.md.

Blind-reconcile cell moves (2026-08-11): #4 O2 4→3, O5 3→4; #6 O5 2→3; #7 O1 4→3, O2 3→4; #8 O5 4→3; #10 O1 2→3. Kept vs blind: #3 O2=4; #9 O1=3, O3=3 (judgment splits).


Anchored scales (1–5) — abstract, outcome-based

Shared band meanings (apply on every axis):

Score Meaning
1 The record contains direct evidence of an adverse effect on this objective.
2 Cited evidence that the instrument fails to move (or undermines) this objective relative to status-quo proxies — not mere absence of research.
3 Mixed / indeterminate, or no project finding either way. Default when evidence is missing.
4 Positive evidence from an analogous, partial, or single-jurisdiction natural experiment / evaluation.
5 Direct multi-setting causal or strong quasi-experimental evidence of a material gain on this objective.

O1 — Realized access

What is scored: timely outpatient appointment or crisis response under the payer that covers the person — not prevalence, not licensed headcount, not HPSA geography alone (ws01, ws02).

1 = Documented deterioration of appointment/crisis access under the relevant payer. 2 = Cited evidence the instrument leaves usable appointment/crisis access unmoved (or worse) on project proxies. 3 = Mixed, proxy-only under KC1, or no finding. 4 = Partial/analogous evidence of improved appointment offer, wait-to-evaluation, or crisis connection under the payer. 5 = Multi-setting evaluation showing material gains in realized appointment or crisis access under the payer.

O2 — Acute and crisis capacity

What is scored: contemporary psych-bed availability (named construct), ED boarding, crisis-continuum throughput, and who can be admitted where (including IMD financing geography). Never 1955%-decline claims (KC2).

1 = Documented worsening of boarding, acute throughput, or admit geography. 2 = Cited evidence the instrument fails to relieve measured contemporary acute/crisis stress. 3 = Mixed, construct-limited, or no finding. 4 = Partial evidence of improved contemporary capacity, boarding, crisis throughput, or admit-financing geography. 5 = Strong quasi-experimental evidence of material acute/crisis capacity gains on locked contemporary constructs.

O3 — Workforce that takes the payer

What is scored: clinicians accepting new patients under Medicaid (and, separately, commercial); safety-net vacancy/participation response — not raw license counts (ws02, ws05).

1 = Documented decline in payer-accepting clinicians or safety-net staffing attributable to the instrument. 2 = Cited evidence the instrument does not open panels / participation under the relevant payer. 3 = Mixed or no finding. 4 = Partial evidence of increased payer participation, service quantity, or safety-net staffing response. 5 = Strong evidence of material participation/staffing gains under the payer.

O4 — Financial protection / parity

What is scored: out-of-pocket exposure, parity compliance in practice, network adequacy that survives a shopper/exam test (ws06). Commercial ERISA majority = band-only / unknown (KC3) unless the cell cites federal EBSA reach.

1 = Documented increase in OOP burden or widening parity/access gaps. 2 = Cited evidence the instrument fails to close measured parity/network gaps in examined plans. 3 = Mixed, KC3 band-only/unknown, or no finding. 4 = Partial evidence of reduced OOP exposure or closed parity/network gaps in examined coverage bands. 5 = Strong evidence of material financial-protection / parity gains spanning the relevant coverage band (including self-funded where claimed).

O5 — State-capacity and federalism load

What is scored: whether an existing agency at the relevant level has demonstrated it can run the instrument; incremental machinery burden (protocol §2).

1 = Documented inability of the responsible agency to operate the instrument (failed implementation on the record). 2 = Cited evidence required machinery exceeds demonstrated agency capacity. 3 = Mixed or no finding. 4 = Closely related instrument already run at material scale by the relevant agency level. 5 = Sustained operation of this instrument class with measured delivery by the responsible agency.


Explicit weight vectors (sum to 1.0)

Protocol §3 named emphases; this scorecard commits the numbers (deviation #17).

Weighting Reader O1 O2 O3 O4 O5 Sum
W1 — Medicaid-SMI enrollee Ongoing care + crisis; Medicaid-dominant SMI band 0.30 0.30 0.20 0.15 0.05 1.00
W2 — Commercially insured parent Outpatient common-disorder care; card that claims parity 0.35 0.05 0.15 0.35 0.10 1.00
W3 — State BH commissioner / Medicaid director Stand up 988, manage IMD, set rates, hold safety-net workforce 0.10 0.20 0.30 0.10 0.30 1.00

Weighted score = Σ (weightᵢ × cellᵢ). Ties reported as ties.


Scores (pass 2 — blind-reconciled)

# Architecture O1 O2 O3 O4 O5 Basis
1 Medicaid behavioral rate floor + administrative simplification 3 3 4 3 4 O3=4: Allegheny MCO rate-change elasticity ≈0.16 plus CCBHC/fee evidence that rates and paperwork bind (H5 Supported) — detectable supply/quantity response, modest and mostly to existing patients (ws05). O1=3: same record shows quantity response ≠ new-panel openings at H2 scale; Bishop/Brahmbhatt acceptance wedge is the problem rate floors aim at, but this pass has no evaluation that a national floor opened new-patient appointments (ws02, ws05). O5=4: states already set Medicaid fee schedules / MCO rates — existing machinery (ws05). O2=3, O4=3: neutral — no finding. Blind matched entire row.
2 CCBHC expansion as default safety-net model 4 3 4 3 4 O1=4 (red-team ↓ from 5; blind matched): Mathematica/ASPE demonstration — adult time-to-eval 9.0→5.4 days (DY1→DY2); clients +~9%; 94% open-access/same-day reported (anchor 9 / H5) (ws05, ws07). Band 4 matches “wait-to-evaluation”; band 5 overshot — metrics are clinic-reported demo before/after, not strong QE, and DY1→DY4 softens (mean 9.1→8.4; within-10-days ~69–73% stable). Proxy under KC1. O2=3 (red-team ↓ from 4; blind matched): crisis is a required CCBHC service, but ED/hospital DID is heterogeneous (not a uniform diversion/throughput win) — mixed on O2 (ws05). O3=4: PPS + certification package moves safety-net capacity/access metrics (ws05). O5=4: §223 demonstration already run; payment-design perishable (GAO-21-104466) (ws05, ws07). O4=3: neutral — evaluation is access/crisis scope, not OOP/parity.
3 IMD repeal or broad MH IMD waiver 3 4 3 3 4 O2=4 (kept vs blind 3): IMD exclusion is a documented Medicaid purchase constraint for adult IMD stays; §1115 SMI/SED patchwork is not national repeal (H3 Supported on statutory/CMS path) (ws03, ws07). O2 band 4 explicitly includes admit-financing geography — that clause carries the 4. Not a proven bed-rebuild: McBain HCRIS finds no significant higher Construct B rates in waiver states (ws03). Blind's evidenced-wash reading is logged in ws09-rescore-log.md. O5=4: states already operate §1115 MH IMD waivers (CRS IF10222 geography) (ws03). O1=3, O3=3, O4=3: neutral — no finding that IMD financing opens outpatient panels or closes commercial parity gaps.
4 Bed rebuild / state hospital reinvestment 3 3 3 3 4 O2=3 (blind ↓ from 4): contemporary Construct A/B need (TAC ~10.8/100k with 52% forensic; NRI 2025 — 90% of states report shortage) is problem documentation, not instrument effect — no evaluation of a bed-rebuild program is in the record (ws03, ws07). H7 Supported demotes asylum-scale as the lead lever; this row is targeted acute/forensic reinvestment pressure, scored honestly as unevidenced on O2. O5=4 (blind ↑ from 3): SMHAs already operate state psychiatric hospitals at material scale with measured reporting (NRI/TAC) — closely related instrument. O1=3, O3=3, O4=3: neutral — no finding.
5 988 + mobile crisis + stabilization continuum completion 4 4 3 3 4 O1=4: crisis-response leg of realized access — national answered/routed volume and answer-rate gains (Vibrant/GAO/KFF; anchor 7) (ws04). Proxy class = contact-center / continuum connection, not outpatient secret-shopper. O2=4: middle-tier continuum evidence (Michigan mobile crisis ~45% lower arrest incidence vs LE-only; BHCC walk-in ↔︎ lower MBD ED use) without large bed rebuilds; 988 itself lacks multi-state causal ED/arrest diversion (H4 Supported) (ws04, ws07). O5=4: 988 transition and state crisis grants already operating (ws04). O3=3, O4=3: neutral — no finding. Blind matched entire row.
6 Parity enforcement with ERISA teeth 3 3 3 4 3 O1=3 (red-team ↓ from 4; blind matched): post-2021 DOL/HHS NQTL network/exclusion corrections are not measured appointment or crisis-connection gains under §2’s secret-shopper/claims proxy class — hold mixed under KC1 (ws06, ws02). O4=4 (examined-band only): CAA reviews found initial comparative analyses insufficient; corrections affected >7.6M participants (H6 Supported) (ws06). Direct federal reach into some ERISA plans — not a census closing KC3 for the self-funded majority. 2024-rule “teeth” perishable — May 2025 nonenforcement of new 2024 provisions (ws06). O5=3 (blind ↑ from 2): EBSA operated the CAA comparative-analysis program at scale and OIG documents investigator capacity ≪ plan census — evidenced wash, not pure below-neutral (ws06). O2=3, O3=3: neutral — no finding.
7 Assisted outpatient treatment expansion 3 4 3 3 4 O1=3 (blind ↓ from 4): NY Kendra’s Law / Swartz et al. measure hospitalization, homelessness, arrest — not appointment offer, wait-to-evaluation, or crisis connection (ws07). Service intensification is a confound. O2=4 (blind ↑ from 3): same Swartz stack — initial 6-month order hospital-admission OR 0.77; renewals stronger — scored as selected-SMI-band acute/throughput evidence (selection is the whole game). Not population bed-capacity. O5=4: NY and CA already run AOT statutes/reporting (ws07). O3=3, O4=3: neutral — no finding. Numerically identical to #3 after reconcile — different reasons; not interchangeable claims.
8 Primary-care behavioral integration / Collaborative Care 3 3 3 3 3 O5=3 (blind ↓ from 4): CoCM codes / FQHC BH expansion exist, but ws05 names same-day billing bans and team payment as state-level binders — wash, not demonstrated scaled delivery (ws05). O1–O4: neutral — no finding — this filing did not land a CoCM access/participation evaluation. All-neutral row, tied with #9 — evidence-density, not endorsement (red-team Attack 4 extended).
9 Workforce liberalization (compacts, supervision ratios, peer specialists w/ Medicaid billing) 3 3 3 3 3 All-neutral evidence-deficit row (judgment splits vs blind 2s on O1/O3). H2 shows licensed ≠ available and warns that headcount-only instruments miss the Medicaid wedge (ws02) — that is scoring guidance, not a measured adverse effect of compact/peer reforms. Blind's own disclosure: no compact/supervision/peer-billing evaluation in the packet. Cells stay at 3 (red-team Attack 4 holds; ws09-rescore-log.md §C).
10 Do-nothing comparator 3 2 2 2 4 O1=3 (blind ↑ from 2): outpatient status-quo proxies are static-poor (Bishop 43.1%; Brahmbhatt 17.8%; Oregon phantoms) and the current 988 path shows measured crisis-connection gains — O1's appointment-or-crisis OR nets to wash (ws02, ws04, phase0). O2=2: IMD default + boarding at Medicaid/uninsured margin persist under patchwork waivers; 988 volume rises without proven diversion (ws03, ws04). O3=2: Medicaid participation wedge remains the flagship finding under current rates/admin (ws02, ws05). O4=2: H6 parity gaps persist; KC3 leaves ERISA-majority commercial O4 band-unknown; 2024-rule nonenforcement (ws06). O5=4: current path is already being run — low incremental machinery (ws04, ws08).

Rankings under three weightings

W1 — Medicaid-SMI enrollee (0.30 / 0.30 / 0.20 / 0.15 / 0.05)

Rank Architecture Weighted score
1 #5 988 + mobile crisis + stabilization 3.65
2 #2 CCBHC expansion 3.55
3 (tie) #3 IMD repeal/broad waiver · #7 AOT 3.35
5 #1 Rate floor + admin simplification 3.25
6 #6 Parity with ERISA teeth 3.15
7 #4 Bed rebuild / state hospital 3.05
8 (tie) #8 Primary-care BH / CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.40

W2 — Commercially insured parent (0.35 / 0.05 / 0.15 / 0.35 / 0.10)

Rank Architecture Weighted score
1 #2 CCBHC expansion 3.60
2 #5 988 + mobile crisis + stabilization 3.50
3 #6 Parity with ERISA teeth 3.35
4 #1 Rate floor + admin 3.25
5 (tie) #3 IMD repeal/broad waiver · #7 AOT 3.15
7 #4 Bed rebuild / state hospital 3.10
8 (tie) #8 Primary-care BH / CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.55

Read carefully: #2’s W2 lead is still driven by O1-proxy safety-net clinic access metrics, not by commercial secret-shopper evidence. #6’s O4=4 is examined-band CAA corrections only — it does not close KC3 for the ERISA-majority commercial parent (red-team Attack 5). Treat #6 as the best examined-band parity instrument on the board, not as resolved commercial O4. AOT’s former W2 near-top is withdrawn (O1 stretch corrected).

W3 — State BH commissioner (0.10 / 0.20 / 0.30 / 0.10 / 0.30)

Rank Architecture Weighted score
1 #2 CCBHC expansion 3.70
2 (tie) #1 Rate floor + admin · #5 Crisis continuum 3.60
4 (tie) #3 IMD repeal/broad waiver · #7 AOT 3.50
6 #4 Bed rebuild / state hospital 3.30
7 #6 Parity with ERISA teeth 3.10
8 (tie) #8 Primary-care BH / CoCM · #9 Workforce liberalization 3.00
10 #10 Do-nothing 2.70

Rank-stability / swing architectures

Stable top pair (survives blind): #5 leads W1; #2 CCBHC leads W2 and W3 (3.60 / 3.70). Both rows were matched cell-for-cell by the blind scorer. Pass-1 “CCBHC leads all three” remains withdrawn (red-team); the post-red-team / post-blind split leadership holds.

Stable bottom: #10 do-nothing is last under every weighting. O1 softened 2→3 (crisis path nets with outpatient failure); O2–O4 status-quo 2s still carry the pole.

Stable near-floor (ignorance, not refutation): #8 and #9 are now identical all-3s (8th–9th everywhere). #8 lost its lone O5=4; #9 judgment splits kept the published neutrals against blind 2s. Rank is evidence-density — not mild endorsement (Attack 4).

H7 demotion holds: #4 bed rebuild never leads. O2 corrected to unevidenced-neutral (need ≠ instrument); O5 rises on SMHA operating capacity. Falls when O1 binds (7th under W1).

Swing — #6 parity with ERISA teeth: O5 wash lifts the row (sole 3rd under W2; 7th under W3). Commercial O4 emphasis still elevates examined-band O4=4; do not read W2 placement as KC3 resolved.

Swing — #7 AOT: O1/O2 swap is W1-neutral and W2-punishing. Former W2 3rd-place near-top withdrawn; now tied with #3 under all three weightings on identical vectors for different reasons.

Near-stable second tier / W1 leader: #5 crisis continuum is 1st / 2nd / tied-2nd. Volume and continuum steelman support it; do not read as a 988 diversion proof (H4).

Flat mid-pack risk: several rows sit between 3.0 and 3.4. #3/#7 identity and #8/#9 identity are the honest flatness finding — reporting distinct ranks there would invent resolution.

Missing row: measurement infrastructure (national by-payer realized-access series) is not among architectures 1–10 — named gap, deviation #19 / Attack 7.


Honesty box — what this scorecard cannot claim

Item Status
KC1 Fires. O1 cells are proxy-scored. There is still no national by-payer wait/acceptance series. The country is arguing about a shortage it does not measure at the point of use (phase0, ws02).
KC1 proxy classes O1 ranks compare unlike proxies (clinic demo waits, Lifeline volume, NQTL corrections, secret-shopper). Do not narrate a single access ladder across rows (red-team Attack 2).
KC2 Fires. No “beds since 1955 / 95% gone” claim enters any cell. #4 uses Construct A/B contemporary evidence only — and O2 is now unevidenced-neutral on instrument effect (ws03).
KC3 Fires. Commercial O4 for ERISA-majority lives is band-only / unknown except where EBSA CAA reviews are cited. State exam evidence does not generalize to self-funded plans (ws06). #6 O4=4 is examined-band only.
H8 Not supported as written. Mental disorders are 12.7% of 2023 disabled-worker awards (MSK 34.0% is the plurality). Stock ≈28.6% is GBMT-12’s object. Do not score architectures on false awards plurality or scarcity→awards overflow (Swenson sign is pathway, not valve) (ws08).
Pass 2 (blind-reconciled) Structurally blinded re-score complete (ws09-rescore-log.md). 7 cells moved; #5/#2 leadership and #10 pole held. Per M10, one re-score is not “final.”
#2 no longer leads all three Post-red-team / post-blind: #5 leads W1; #2 leads W2/W3. Pass-1 all-three CCBHC headline withdrawn.
#7 W2 near-top withdrawn Hospitalization OR is not O1; AOT falls out of W2 top tier after reconcile.
#2’s W2 lead Remains an O1-proxy lead from Medicaid safety-net evaluations, not proof CCBHCs fix commercially insured outpatient parity.
#6 ≠ KC3 closed Examined-plan CAA corrections ≠ ERISA-majority census. W2 3rd place is narrowed, not a commercial O4 win.
#1 does not “solve H2” Rate floors + admin simplification are directionally correct and elasticity-limited; Allegheny response was mostly existing patients (ws05).
#3 ≠ bed rebuild IMD repeal/waiver scores on financing geography, not on McBain-proven bed growth (null so far) (ws03).
#4 O2 ≠ contemporary need Need documentation stays in the workstreams; it does not score the rebuild instrument above 3.
#7 ≠ system fix AOT evidence is selected-band acute/engagement; selection/regression-to-mean footnotes are mandatory (ws07).
Neutral cells are not failures — and mid-pack is not endorsement Held at 3 for missing evidence (especially #8/#9 all-neutral; many O4 Medicaid-architecture cells). Inventing 1s/2s from silence would repeat the childcare / pre-amendment elder-care failure mode. Conversely, #8/#9 placement is ignorance density, not mild support (Attack 4).
Measurement architecture missing Protocol KC1 makes measurement the headline shape; architectures 1–10 omit a scored “publish national by-payer realized-access series” row. Whitepaper / next pass must score it or record explicit exclusion (Attack 7 / deviation #19). VA access standards are the domestic measurement precedent (ws07).

Cells held at neutral (3) for missing evidence (non-exhaustive)


Freeze / next passes

Protocol architectures and O1–O5 objectives unchanged. Weight numbers and symmetric below-neutral rule logged as deviations #17–#18. Red-team amendments logged as deviation #20. Blind reconcile logged as deviation #21. Next owed under M6/M10: whitepaper synthesis may proceed on this twice-checked board with ranks still provisional under M10; measurement row still owed (deviation #19). Verification Protocol Phase 1 not yet run on this filing.

← All Mental health research documents Sources digest Read the whitepaper