GUBMENTPlain talk · policy frontier
Filings / Mental health / Sources / Blind re-score — raw scorer output (GB
GBMT-11 · Research record · No. 11

Blind re-score — raw scorer output (GBMT-11)

mental-health/research/ws09-blind-scores-2026-08-11.md
This is a working research document from the mental health filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 11, not a summary of it.

Date: 2026-08-11 Model: claude-opus-5 via the Anthropic Message Batches API (scripts/batch-rescore.py, batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx). Blinding: structural. The scorer received one tool-less request containing the scoring brief (ws09-blind-brief-2026-08-11.md), the anchored scale (00-scales.md), and the evidence packet below inlined as text. It had no filesystem access and no search, so it could not reach ws09-scorecard.md, ws09-red-team-log.md, or deviations-log.md by any route.

Reconciled in: ws09-rescore-log.md. What follows after the packet manifest is the scorer's output, verbatim and unedited.


Packet manifest

Evidence was inlined from /private/tmp/gubment-batch-rescore-gbmt11-12/mental-health/ (batch working directory). Evidence files are not re-copied into git — they match the committed workstreams/phase0 on origin/main (SHA-256 checked at reconcile). Brief + raw scores are the research artifacts.

Path in packet Bytes Repo counterpart
BRIEF.md 3784 ws09-blind-brief-2026-08-11.md
SCORES.md 25266 this file (body below)
evidence/00-scales.md 4126 packet-only extract of scorecard anchors (not a separate repo file)
evidence/phase0-findings.md 12108 mental-health/docs/phase0-findings.md
evidence/ws01-baseline.md 8335 mental-health/research/ws01-baseline.md
evidence/ws02-licensed-available.md 8460 mental-health/research/ws02-licensed-available.md
evidence/ws03-imd-acute.md 9561 mental-health/research/ws03-imd-acute.md
evidence/ws04-crisis-988.md 5970 mental-health/research/ws04-crisis-988.md
evidence/ws05-medicaid-rates.md 7914 mental-health/research/ws05-medicaid-rates.md
evidence/ws06-parity.md 8812 mental-health/research/ws06-parity.md
evidence/ws07-precedents.md 14126 mental-health/research/ws07-precedents.md
evidence/ws08-political-economy-seam.md 10394 mental-health/research/ws08-political-economy-seam.md

What the scorer did not receive: ws09-scorecard.md, ws09-red-team-log.md, deviations-log.md, any site/mental-health/*, and this reconciliation log.

Evidence SHA note: At reconcile (2026-08-11), every evidence/ws0*.md and phase0-findings.md byte-matched the corresponding files on this branch (origin/main tip). 00-scales.md is the authored abstract scale shipped in the brief packet (aligned with the scorecard's O1–O5 anchors, including the symmetric evidence floor already baked into pass 1).


GBMT-11 Mental Health — blind re-score

1. The matrix

O1_realized_access O2_acute_crisis_capacity O3_workforce_takes_payer O4_financial_protection_parity O5_state_capacity_federalism
a1_medicaid_rate_floor 3 3 4 3 4
a2_ccbhc_expansion 4 3 4 3 4
a3_imd_repeal_waiver 3 3 3 3 4
a4_bed_rebuild 3 3 3 3 4
a5_crisis_continuum 4 4 3 3 4
a6_parity_erisa 3 3 3 4 3
a7_aot_expansion 3 4 3 3 4
a8_collaborative_care 3 3 3 3 3
a9_workforce_liberalization 2 3 2 3 3
a10_do_nothing 3 2 2 2 4

Cell IDs below: rows a1–a10 = 1–10, columns O1–O5 = A–E.

2. Cell-by-cell basis

a1 — Medicaid rate floor + administrative simplification

Cell Score Basis Flag
1A 3 ws05: Allegheny MCO rate-change design finds elasticity ≈0.16 (20% rate rise → ~3.2% more services) but "mostly to existing patients rather than new panel openings"; Zhu Health Affairs 2022 cross-section null (state Medicaid-to-Medicare indices uncorrelated with participating-psychiatrist supply); Decker null on reported participation. No wait/appointment-access outcome measured for any rate change. Real evidence both directions, none of it on appointment access. evidenced-wash
1B 3 Nothing in the record links Medicaid outpatient rate/prior-auth changes to boarding, bed supply, or crisis throughput. unevidenced-neutral
1C 4 ws05: Allegheny within-market staggered code-level rate increases → detectable increase in service quantity (O3 anchor 4 names "service quantity"); H5 adjudicated Supported ("rates and paperwork bind; elasticities modest"); Cunningham/CMS MMRR — rates + payment delay + billing requirements co-cited by physicians limiting Medicaid; NACHC same-day-billing bans as binding. Countervailing nulls (Zhu cross-section, Decker) keep this off 5. judgment-call (anchor's "service quantity" wording is doing the work; the instrument's own stated goal — panels opening — is not evidenced)
1D 3 No OOP/parity finding for Medicaid rate policy in the record. unevidenced-neutral
1E 4 ws05: state Medicaid FFS fee schedules for a common psychiatrist basket exist and are measurable in every state (indices 0.46 PA → 2.34 NE); an MCO in Allegheny actually executed staggered rate changes across many codes. The relevant agency level demonstrably runs rate-setting at scale. Admin-simplification half is less demonstrated (same-day billing bans persist). evidenced

a2 — CCBHC expansion

Cell Score Basis Flag
2A 4 ws05/anchor 9 (ASPE/Mathematica §223 evaluation): adult mean time to initial evaluation 9.0 → 5.4 days DY1→DY2; clients served +~9% across seven states (PA +23%); 94% of CCBHCs report open-access/same-day. Not 5: clinic-reported, pre/post without comparison group, and DY1→DY4 flattens (within-10-days stable ~69–73%; mean 9.1 → 8.4 days). evidenced
2B 3 ws05: claims DID heterogeneous — PA −13% BH-related ED visit count (no change in any-ED probability), MO +5.7% BH ambulatory visits, "hospitalizations/ED often null"; ws05 states explicitly this is "state-heterogeneous, not a uniform national win." I honoured the record's own mixed characterization rather than picking the PA positive. evidenced-wash
2C 4 ws05: PPS + certification package associated with more clients served (304,988 → 332,135 across reporting states) and near-universal open-access scheduling — a safety-net service-quantity/capacity response under Medicaid. evidenced
2D 3 No OOP or parity finding for CCBHCs in the record. unevidenced-neutral
2E 4 §223 demonstration operated by state Medicaid agencies with CMS across DY1–DY4 with measured delivery reporting (Mathematica/ASPE/RTC); >60 dual CHC–CCBHC certifications (NACHC). Held at 4 rather than 5 because GAO-21-104466 documents mixed spending effects and CMS guidance gaps on aligning PPS rates with costs / avoiding duplication — payment design is load-bearing and perishable. judgment-call (4 vs 5 was close on a literal reading of the O5 anchor)

a3 — IMD repeal / broad MH IMD waiver

Cell Score Basis Flag
3A 3 Nothing in the record on outpatient appointment or crisis access effects of IMD waivers. unevidenced-neutral
3B 3 ws03/ws07 both directions: (+) statute + CMS SMDL #18-011 + CRS IF10222 document that FFP for adult 21–64 IMD stays is barred by default, so the waiver directly changes admit-financing geography — a documented purchase constraint being relieved; (−) McBain et al. PLOS Medicine 2025 (HCRIS): MH IMD §1115 waiver states show no statistically significant higher psychiatric bed rates. ws03 explicitly says the null is "a caution for architecture scoring, not a full refute of the financing constraint" and "do not overclaim that waivers have already rebuilt Construct B supply." evidenced-wash
3C 3 No finding on payer-accepting clinicians. unevidenced-neutral
3D 3 No OOP/parity finding. unevidenced-neutral
3E 4 CRS IF10222 (as of 2025-01-14): ~15 states + DC with approved MH IMD §1115 authority plus a pending queue; CMS runs the SMI/SED opportunity. The waiver machinery is demonstrably operated at material scale by CMS + state Medicaid agencies. Note: the repeal half of this row requires Congress, and the record contains no evidence on statutory repeal capacity. evidenced

a4 — Bed rebuild / state-hospital reinvestment

Cell Score Basis Flag
4A 3 No record finding linking inpatient reinvestment to outpatient/crisis appointment access. unevidenced-neutral
4B 3 This is where the evidence floor bites hardest and I want it visible. The record documents need, not instrument effect: NRI 2025 (90% of responding states report inpatient psych bed shortage; forensic most cited), TAC 2023 Construct A ~36,150 / 10.8 per 100k with 52% forensic occupancy, McBain Construct B stable ~28.4/100k with STACH psych-unit share 24.9% → 20.6%, Pearlmutter multi-site boarding (admitted median 5.63h; Medicaid/uninsured >2× odds of ≥24h). Per discipline 2, problem size is not evidence of an instrument's effect, and no evaluation of a bed-rebuild or state-hospital reinvestment program appears in this base. The Australia low-bed/high-readmission pattern (ws07 counter-steelman) is a cross-national correlation, not an instrument evaluation. I did not dock this cell for H7's "wrong primary architecture" verdict, which is a ranking statement, not a within-cell failure finding. unevidenced-neutral (deliberate: I reasoned about the mechanism and refused to score it)
4C 3 No finding on payer-accepting clinicians or state-hospital staffing response to reinvestment (NRI reports shortages, not an instrument effect). unevidenced-neutral
4D 3 No OOP/parity finding. unevidenced-neutral
4E 4 State mental health agencies demonstrably operate state psychiatric hospitals at material scale with measured use reporting (NRI Use of State Psychiatric Hospitals, 2025 Profiles; 43/48 states responding; TAC staffed-bed censuses). Held at 4 not 5 because the instrument as specified is expansion/reinvestment, and the record contains no measured delivery of a capacity-expansion program. judgment-call

a5 — 988 + mobile crisis + stabilization continuum

Cell Score Basis Flag
5A 4 ws04: answered network contacts (ex-VCL) 161,267 Jan 2022 → 403,544 Jan 2024 (~2.5×); GAO-26-108114 ~19.1M contacts routed Jul 2022–Sep 2025, calls +87%, overall ~+90%; national answer rate 70% (May 2022) → 89% (May 2024), wait 2:20 → 1:31. This is measured improvement in crisis connection, which O1 explicitly admits. Not 5: operational before/after series with no counterfactual design, Lifeline covers ~200 of ~544 crisis centers, and ws04 warns answer-rate gains measure contact-center throughput. Also Swanson (MI) mobile-crisis response as the field-response leg. evidenced
5B 4 ws04/ws07: Kalb et al. Health Serv Res 2025 — each additional zip-level walk-in behavioural health crisis centre associated with ~2.8% lower mean mental/behavioural-disorder ED utilisation; Swanson et al. Michigan multi-site mobile crisis (IPTW) — ~45% lower 11-month arrest incidence vs law-enforcement-only. Both single-jurisdiction/associational → 4, not 5. I honoured ws04's instruction to keep 988 volume separate from diversion: the 988-specific diversion claim is unproven, and absence of a multi-state 988 causal estimate is not scored as failure. evidenced
5C 3 No finding on crisis-workforce payer participation or safety-net staffing response. unevidenced-neutral
5D 3 No OOP/parity finding for crisis services. unevidenced-neutral
5E 4 GAO-26-108114 + Vibrant KPIs: SAMHSA/administrator + state centres have run 988 routing continuously Jul 2022–Sep 2025 with published delivery metrics (~76% local answer for calls); Arizona comprehensive crisis system as a state-run continuum instance. Held at 4 not 5 because the row is continuum completion — the stabilization/mobile tier is not demonstrated at national scale — and GAO notes state-vs-administrator metric divergence and unmet text/chat local-answer goals. judgment-call

a6 — Parity enforcement with ERISA teeth

Cell Score Basis Flag
6A 3 Both directions, neither on appointment access: (+) EBSA corrections since Feb 2021 removed exclusions and fixed network-monitoring practices benefiting >7.6M participants in >72k plans (2024 MHPAEA RTC); NY 11 NYCRR 38 sets numeric behavioural wait/network standards; (−) comparative analyses were effectively all insufficient on initial submission (2022/2023 RTCs) and GAO-22-104597 documents continuing access challenges. Under KC1 there is no realized-access series to test against. evidenced-wash
6B 3 No acute/crisis capacity finding for parity enforcement. Wit is confined to holdings and is class-procedure/plan-interpretation, not capacity. unevidenced-neutral
6C 3 "Fixed network-monitoring practices" is a compliance action, not a measured change in clinicians accepting commercial coverage. No finding. unevidenced-neutral
6D 4 ws06: post-CAA-2021 federal comparative-analysis program produced final noncompliance determinations and corrections (exclusion removals, network-monitoring fixes) across >72k plans / >7.6M participants, including self-funded plans (FY2023 EBSA/CMS fact sheet cites self-funded plans with MH/SUD out-of-network use far above M/S); plus independent state exam root (MA c.26 §8K, 21 carriers; NY §343 biennial reports) finding fully insured gaps. Capped at 4 per KC3: EBSA reach is investigative spot-check, not census (2.6M plans / ~136M participants; DOL OIG 09-25-001: <1 investigator per ~16,472 plans), and the May 2025 statement non-enforces the new 2024-rule provisions. evidenced
6E 3 Both directions: (+) EBSA has actually operated the CAA comparative-analysis program at scale and MA/NY show state exam machinery working on fully insured plans; (−) DOL OIG 09-25-001 documents investigator capacity far below the plan universe and loss of supplemental NQTL funding; GAO-23-105642 notes no federal numeric network-adequacy standard; May 2025 non-enforcement of the 2024 rule. The census-style reach the row's spec assumes exceeds demonstrated capacity. evidenced-wash

a7 — Assisted outpatient treatment

Cell Score Basis Flag
7A 3 The NY/CA evaluations measure hospitalization, homelessness and arrest — not appointment offer, wait-to-evaluation, or crisis connection. ws07 notes service intensification (ICM/ACT) accompanies orders, but presents that as a confound, not a measured access outcome. There is also no evidence bearing on population outpatient access (the row spec's own warning). evidenced-wash (I considered 4 on the intensification reading and rejected it as mechanism)
7B 4 ws07: Swartz et al. Psychiatric Services 2010 (NYS OMH + Medicaid claims linkage, 3,576 AOT consumers 1999–2007): initial 6-month order → hospital admission OR 0.77, renewal → OR 0.59, hospital-days odds ~0.80–0.84; Laura's Law DHCS county-reported drops in hospitalization/LE contact among enrollees corroborate direction. ws07 calls this "the best domestic identification stack" and rates confidence "strong on … NY AOT administrative evaluations." Selection discipline applied: this is a selected-SMI band (prior hospitalization/non-adherence criteria), pre/post, with regression-to-mean and co-delivered ACT as live confounds — which is why it is 4 and not 5, and why it does not travel to population acute capacity. evidenced (with selection caveat)
7C 3 No finding on payer participation. unevidenced-neutral
7D 3 No OOP/parity finding. unevidenced-neutral
7E 4 NY OMH has operated Kendra's Law since 1999 with statewide administrative/Medicaid-linked reporting; CA DHCS produces legislative reports on Laura's Law. Held at 4 not 5 because ws07 documents county opt-out and reporting inconsistency in CA, so the class is not uniformly operated with measured delivery. judgment-call

a8 — Collaborative Care / primary-care integration

Cell Score Basis Flag
8A 3 This evidence base contains no CoCM or integration outcome evaluation at all — no IMPACT-class trial, no wait or appointment effect. ws05's only material on architecture #8 is NACHC on financing barriers (same-day billing bans, inadequate team-based BH reimbursement, documentation burden), which documents adoption friction, not that adoption fails to improve access. Floor applies upward: I did not credit the mechanism. unevidenced-neutral
8B 3 No finding. unevidenced-neutral
8C 3 NACHC describes health centres as "a major BH delivery surface" — a scale descriptor, not an effect. Per discipline 2 that does not lift the cell. unevidenced-neutral
8D 3 No finding. unevidenced-neutral
8E 3 Both directions: (+) CHCs deliver BH at material scale with >60 dual CHC–CCBHC certifications and telehealth/CMHC partnerships; (−) ws05's explicit scorecard guidance is that #8 "hits same-day billing and team payment as state-level binders," i.e. some states' Medicaid rules currently prevent billing the integrated encounter. Nets to mixed. evidenced-wash (I was genuinely torn between 3 and 4 here)

a9 — Workforce liberalization

Cell Score Basis Flag
9A 2 ws02 documents failure of the exact causal channel this instrument relies on: licensed/directory supply does not become usable access. Bishop JAMA Psychiatry 2014 — psychiatrists' Medicaid acceptance 43.1% while licensed supply exists; Brahmbhatt & Schpero JAMA 2024 — only 17.8% of directory-listed Medicaid prescribing clinicians in 4 cities were reachable, accepting and offering an appointment; Zhu Health Affairs 2022 — 58.2% of Oregon Medicaid directory listings phantom (67.4% MH prescribers); HRSA MH HPSA designated population rose 122M → 137M → 157.1M with no corresponding realized-access movement, and ws02 rules HPSA counts "licensed/assigned supply geography, not payer acceptance." ws02's Implications for scorecard instructs directly: "Architectures that only grow licensed headcount without rate/network/directory integrity score poorly on O1/O3" — honoured per discipline 6. evidenced, with disclosure: no compact / supervision-ratio / peer-billing instrument is itself evaluated anywhere in this base, and the peer-specialist-Medicaid-billing sub-instrument is a payer-participation lever the headcount critique does not reach. A scorer who treats the ws02 finding as mediating-channel evidence rather than instrument evidence would put this at 3.
9B 3 No finding on crisis/acute capacity from workforce liberalization. unevidenced-neutral
9C 2 Same citations as 9A; O3 is the axis ws02's instruction names explicitly. Bishop is the record's "existence proof that payer participation can dominate headcount" (ws02 §3 baseline non-substitution #2). evidenced (same disclosure as 9A)
9D 3 No finding. unevidenced-neutral
9E 3 Nothing in the record on compact commissions, state licensing-board capacity, or peer-certification machinery. unevidenced-neutral

a10 — Do-nothing comparator

Cell Score Basis Flag
10A 3 Split, and I am reporting the split rather than resolving it: (−) on outpatient, the status quo's measured proxies are static-poor — Brahmbhatt 17.8% usable access with median waits to 64 days (LA, IQR to 126), Bishop 43.1%, Oregon phantom rates, ~48% of AMI untreated; (+) on crisis, the current 988 funding path is exactly where the measured connection gains happened (answer rate 70% → 89%, wait 2:20 → 1:31, ~+90% routed contacts). O1 covers appointment or crisis response, so these net. evidenced-wash
10B 2 Cited evidence the status quo does not relieve measured contemporary acute stress: NRI 2025 — 90% of responding states report inpatient psych bed shortage (37 forensic, 38 non-state); TAC 2023 Construct A 10.8/100k with 52% forensic occupancy; McBain — Construct B flat 28.1 → 28.4/100k while STACH psych-unit share fell 24.9% → 20.6% and 846 counties (~244M residents) saw IPB rate declines; Pearlmutter multi-site boarding with Medicaid/uninsured >2× odds of ≥24h ED stay; patchwork IMD waivers leave adult IMD purchase barred by default outside ~15 states. Not 1 because the national contemporary bed rate is stable, not declining. evidenced
10C 2 Bishop 43.1% Medicaid acceptance (vs 73.0% other physicians), Brahmbhatt 17.8% usable, Zhu phantom 58.2%/67.4% — panels are documented as not open under the status quo, on two independent methods (ws02 H2 Supported). evidenced
10D 2 ws06 states it directly: "Do-nothing / parity-on-paper (architecture #10) fails H6's evidence bar." Supporting: 2022/2023 RTCs found reviewed comparative analyses insufficient on initial submission; FY2023 enforcement continues to find network-composition and exclusion NQTL problems; KC3 fires — state exams cannot reach the self-funded ERISA majority, and there is no federal numeric network-adequacy standard (GAO-23-105642). evidenced
10E 4 No incremental machinery, and the constituent instruments are demonstrably operating: GAO-26-108114 (~19.1M 988 contacts routed over 39 months), CRS IF10222 (approved MH IMD waivers in ~15 states + DC), EBSA operating MHPAEA jurisdiction over ~2.6M plans / ~136M participants with corrections since Feb 2021. Per discipline 5 I applied no comparator penalty, and this is the honest consequence of the O5 axis: doing nothing scores well on "can the agency run it." evidenced (and flagged as an axis artifact, not a merit finding)

3. What my matrix implies

On my numbers the board is mostly flat, and the flatness is the finding: 32 of 50 cells sit at 3, and 15 of those are unevidenced-neutral. a2_ccbhc_expansion and a5_crisis_continuum tie as strongest — they are the only rows with cited positive effect evidence on two or more substantive axes (CCBHC: measured wait-to-evaluation and client-volume gains under Medicaid; crisis continuum: measured crisis connection plus two independent continuum studies on ED use and arrests). a9_workforce_liberalization and a10_do_nothing tie as weakest, for opposite reasons: a9 because the record documents failure of the licensure→usable-access channel it depends on, a10 because the record documents unrelieved acute stress, closed panels and unclosed parity gaps on the status-quo trajectory. a3_imd_repeal_waiver and a4_bed_rebuild come out numerically identical (3,3,3,3,4) — I am saying that plainly: the evidence base gives one a documentary financing constraint offset by a bed-count null, and the other essentially nothing but need documentation, and both land on the neutral floor. a1 and a7 sit mid-board with two evidenced cells each; a6 is single-cell evidenced (O4) and constrained by KC3; a8 is the only row with zero evidenced non-neutral cells.

4. Evidence-coverage count

Architecture Cells scored off 3 on cited evidence Cells held at 3 as evidenced-wash Cells held at 3 as unevidenced-neutral
a1_medicaid_rate_floor 2 (O3, O5) 1 (O1) 2
a2_ccbhc_expansion 3 (O1, O3, O5) 1 (O2) 1
a3_imd_repeal_waiver 1 (O5) 1 (O2) 3
a4_bed_rebuild 1 (O5) 0 4
a5_crisis_continuum 3 (O1, O2, O5) 0 2
a6_parity_erisa 1 (O4) 2 (O1, O5) 2
a7_aot_expansion 2 (O2, O5) 1 (O1) 2
a8_collaborative_care 0 1 (O5) 4
a9_workforce_liberalization 2 (O1, O3) 0 3
a10_do_nothing 4 (O2, O3, O4, O5) 1 (O1) 0

Total: 18 cells scored on cited evidence, 8 evidenced-wash, 24 unevidenced-neutral (one row, a2, has its 3s split 1/1).

5. What I could not score, and what would change my mind

Structurally unscoreable in this base:

What would move specific cells:

One anchor-fit complaint. O5's anchors reward instruments that are already running, which mechanically advantages the do-nothing comparator (10E = 4) and disadvantages nothing except genuinely novel machinery. I applied the anchors as written rather than adjusting, but the resulting O5 column is nearly constant (nine 4s and one 3 out of ten... in fact seven 4s and three 3s) and carries little discriminating information on this board.

← All Mental health research documents Sources digest Read the whitepaper