Date: 2026-08-03 (first pass); red-teamed 2026-08-03 (red-team-log.md); independently re-scored and reconciled 2026-08-04 (ws10-rescore-log.md) — two headline corrections (the supply-side row's split verdict; the prevention row re-anchored to the record's evidence about the instrument rather than the size of its target), ~30 cells reconciled, one honestly-disclosed partial contamination on A3's rank noted in the log. This is the v3 (twice-checked) matrix below; superseded cells are visible in git history.
Method
11 candidate architectures (10 from the protocol's §10 seed list + the status-quo-drift comparator required by M7) scored on 8 anchored dimensions, each cell citing the workstream finding it rests on. Scored 1 (weak) – 3 (strong) per dimension; dimension 3 (fiscal cost) and dimension 8 (speed) are inverted so 3 = cheap/fast, consistent with "higher is better" throughout. Ranked under four explicit objective weightings per §1.5's objective-function decision: mortality-first, order-first (public order/political durability), liberty-first, and equal weights. Rank stability across weightings — not any single ranking — is the headline finding, per the GBMT-1/GBMT-2 precedent that this is the most useful thing a scorecard can report.
Dimensions (anchored)
| # | Dimension | 1 (weak) | 2 (moderate) | 3 (strong) |
|---|---|---|---|---|
| D1 | Evidence strength for core mechanism | Single-root, contested, or no controlled evidence | Directional/observational support, real caveats | RCT or multi-study peer-reviewed convergence |
| D2 | Legal executability | Requires new statute (Congress) | Mixed — partial agency action possible | Achievable via existing agency rulemaking alone |
| D3 | Fiscal cost (inverted: 3=cheap) | Requires major new appropriation | Moderate cost, existing funding streams stretchable | Near-zero net new cost |
| D4 | Mortality-impact potential | No credible mechanism to overdose deaths specifically | Plausible but unquantified/indirect | Direct, evidenced mechanism to overdose mortality |
| D5 | CJ-footprint reduction potential | Expands or is neutral to enforcement/incarceration footprint | Modest reduction | Directly reduces arrests/incarceration exposure |
| D6 | Political durability / public-order risk (3=durable) | High visible-disorder backlash risk (per §8 pattern) | Moderate risk, contested locally | Low backlash risk, broad or bipartisan constituency |
| D7 | Liberty impact (3=least restrictive) | Expands state coercive power over individuals | Neutral | Expands individual autonomy/reduces coercion |
| D8 | Speed to effect (inverted: 3=fast) | Multi-year implementation lag likely | 1–2 year lag | Effects observable within ~1 year |
Scorecard
| Architecture | D1 | D2 | D3 | D4 | D5 | D6 | D7 | D8 | Basis (workstream) |
|---|---|---|---|---|---|---|---|---|---|
| A1. Methadone deregulation (pharmacy/office dispensing) | 2 | 1 | 3 | 2 | 1 | 2 | 3 | 1 | §4: mixed international retention evidence, but France (§Phase0) and Australia pharmacy models show no mortality/diversion increase. §6: OTP monopoly is statutory (21 U.S.C. §823(h)) — Congress required, MOTAA pending |
| A2. MOUD-everywhere (ED-initiation, jail/prison mandates, telehealth permanence) | 2 | 2 | 2 | 2 | 1 | 3 | 2 | 2 | §4: X-waiver repeal natural experiment shows legal access alone doesn't move patient volume (provider willingness binds) — tempers this architecture's D4/D1 score despite strong underlying MOUD efficacy evidence |
| A3. Contingency-management legalization (safe-harbor expansion) | 3 | 3 | 3 | 2 | 1 | 3 | 2 | 3 | §4: strongest evidence base for stimulant UD specifically (43.1% of deaths are opioid+stimulant polysubstance per §3) — OIG concurs (85 FR 77791). §6: OIG has clear standing rulemaking authority (Pub. L. 100-93 §14(a)). BASIS CORRECTED 2026-08-10: |
| A4. Treatment-industry quality regime (outcome measurement + conditional funding) | 1 | 2 | 2 | 2 | 1 | 2 | 2 | 1 | §4: no federal outcome-reporting requirement exists today; EKRA addresses fraud, not quality. No controlled evidence located that outcome-conditioning improves results — genuinely untested architecture, scored low on D1 for that reason, not dismissed |
| A5. Harm-reduction scale-up (naloxone saturation, drug checking, SCS at evidenced scale) | 2 | 2 | 2 | 3 | 2 | 1 | 3 | 3 | §3: naloxone/FTS legal-access fight is largely already won (OTC 2023, 45-state FTS decriminalization); causal contribution to the 2023-25 decline is asserted, not quantitatively isolated. §8: highest public-order/backlash risk of any architecture — the SF/Oregon reversal pattern applies most directly here |
| A6. Decriminalization + dissuasion + prompt treatment funding (Portugal-style, explicitly not Oregon-style) | 1 | 1 | 1 | 2 | 3 | 1 | 3 | 2 | §9 (Portugal, Oregon): both precedents show decrim-without-prompt-treatment-funding underperforms; the >1-year funding rollout delay is the common failure mode in both cases. §8: this is the architecture most exposed to the documented backlash pattern (both flagship real-world attempts reversed within 2-4 years) |
| A7. Cannabis federal redesign (descheduling + state opt-in) | 2 | 2 | 3 | 1 | 2 | 2 | 3 | 2 | §6: narrow Schedule III move already done via rulemaking (April 2026) but full descheduling/banking needs Congress and is actively litigated. §7: CA/WA illicit markets persist at 50-60% even post-legalization — legalization alone doesn't solve market design, tempering D1 |
| A8. Supply-side modernization (precursor diplomacy, interdiction) | 2 | 3 | 2 | 2 | 1 | 3 | 1 | 2 | §5: 40-year price/purity paradox is the single strongest piece of evidence against this mechanism (H5.1 supported) — real prices fell ~80% through the exact decades of maximal enforcement spending escalation. Scored low on D1/D4 accordingly, though the 2026 fentanyl-supply-shock paper (§3) is a partial, contested counter-data-point |
| A9. Prevention/tobacco-playbook transplant (taxation, age-gating, marketing limits, incl. alcohol) | 2 | 2 | 3 | 1 | 1 | 2 | 1 | 1 | §2: alcohol (~178k/yr) and tobacco (~450-480k/yr) deaths dwarf peak overdose deaths (~112-114k) — this architecture targets the numerically larger problem with the field's best-evidenced policy toolkit (tobacco control's ~50%-decline record). §7: the alcohol industry's lobbying scale ($541M 1998-2020) and its documented pivot to shaping cannabis's rules is the clearest evidence this architecture faces organized, well-resourced opposition once it targets alcohol specifically |
| A10. Enforcement steelman (focused deterrence + drug courts, explicitly not HOPE-style) | 3 | 3 | 2 | 2 | 2 | 3 | 1 | 2 | §5: focused deterrence has real RCT support (~16-23% crime reduction) for gang/gun violence, not drug-market disruption; drug courts work for serious/dependent offenders specifically (38% vs 50% recidivism). HOPE-style supervision excluded from this architecture on the strength of its failed 4-site replication |
| A11. Status quo drift (comparator, per M7) | — | 3 | 3 | 2 | 2 | 3 | 2 | 3 | §3: overdose deaths are already falling sharply (~37-38% peak-to-2025) on the current trajectory, for reasons (§3's supply-shock/exposure-depletion findings) largely independent of any single policy architecture. Every architecture above must beat this counterfactual, not zero |
Rankings, re-scored (full reconciliation: ws10-rescore-log.md)
Reconciled sums (D1–D8): A3 20 · A5 18 · A10 18 · A11 18-of-7-dimensions (per-dimension mean 2.57, edging A3's 2.50 — the comparator's strength is partly that deaths are already falling for reasons no architecture here caused, which is §3's own headline applied to this table) · A7 17 · A2 16 · A8 16 · A1 15 · A6 14 · A4 13 · A9 13.
A3 (contingency-management legalization) remains the top-ranked architecture and the filing's clearest first move — with the re-score's contamination caveat stated in the log rather than hidden: the protocol's status line leaked A3's headline to the second scorer, so top-rank agreement is discounted; A3's per-cell case (trial evidence, live OIG rulemaking authority, near-zero cost, demonstrated one-year speed) was independently re-derived and stands.
Two v1 claims are withdrawn on reconciliation:
- "A8 ranks last under every weighting" — the supply-side row now carries a split verdict: the price/purity paradox refutes the interdiction leg, but the record's own §3 attributes the historic 2023–25 mortality decline to a supply shock, the filing's best observed mortality correlation. Scoring the architecture 1-on-mortality while Part 3 credits a supply-side mechanism for the decline was internally inconsistent; the reconciled row (Σ16) reflects the split, with the caveats (fentanyl price direction, West Coast timing) named.
- A9's second-place rank — the prevention row was scored on the size of its target rather than the record's evidence about the instrument: D4's anchor reads "overdose deaths specifically," no workstream ever executed the §9 tobacco arm, and the tobacco model's absence-of-arrests is not a reduction of the existing drug-arrest footprint. A9 falls to bottom-tier (Σ13) — with the rubric note, carried honestly, that the overdose-specific mortality anchor structurally penalizes the architecture aimed at the numerically larger legal-drug body count. That is a scoring-frame limitation, not a verdict that prevention is a bad idea.
The bottom is now A4 and A9 (13 each) — the genuinely-untested idea and the out-of-frame one — not A8. The largest weighting-dependent swings remain A6 (best CJ-footprint and liberty case, worst durability case per §8's documented reversal pattern) and A10 (strong order/executability, weakest liberty) — both should still be presented with their weighting-dependence explicit.
Phase 1 verification effect on this scorecard (2026-08-10)
The Phase 1 fact-check (verification-log.md) did not re-score any cell — re-scoring belongs to a fresh, structurally blinded pass, and patching cells here would bypass the reconciliation M6 requires (the same reasoning deviations #19 and #20 already applied). It did establish that three cells now rest on bases that primary text contradicts, and those are recorded here as owed rather than silently left standing:
| Cell | Current | Why it is now in question |
|---|---|---|
| A3 D2 (legal executability = 3, "achievable via existing agency rulemaking alone") | 3 | Still rulemaking-only, but the anti-kickback statute remains and a safe harbor only shelters conduct from it; the Secretary must act in consultation with the Attorney General |
| A3 D8 (speed = 3, "effects observable within ~1 year") | 3 | RIN 0936-AA13's NPRM target is July 2027; nothing is pending |
| A1 D2 (legal executability = 1, "requires new statute") | 1 | §823(h) requires a separate registration, not OTP exclusivity — a materially regulatory component the "Congress required" score did not account for |
Two of the three move against the filing's headline and one moves for a rival architecture, which is the shape a real check produces.
Phase 2 steelman effect on this scorecard (2026-08-10)
The Phase 2 steelman (steelman-log.md) also did not re-score any cell, for the same reason. It published a sensitivity analysis — the arithmetic of this same reconciled matrix under three explicitly stated readings — and it recorded three further cells as contested. All of it is owed to the blinded re-score, not settled here.
| Cell | Current | Why it is now in question |
|---|---|---|
| A3 D3 (fiscal cost = 3, "near-zero net new cost") | 3 | Prices the rule change, not the service. What states actually stand up is a per-site CM Coordinator, 36 point-of-care tests per patient, an incentive-management platform, a pre-launch readiness review and ongoing fidelity monitoring (§4) |
| A3 D4 (mortality impact = 2, "plausible but unquantified") | 2 | Moves in the filing's favour. Coughlin et al., Am J Psychiatry 2025;182(11):1016–1023 — first real-world evidence, 41% lower one-year all-cause mortality (aHR 0.59, 95% CI 0.36–0.95) in a matched VHA cohort |
| A9 D1 and D4 (evidence = 2, mortality = 1) | 2, 1 | D4's "overdose deaths specifically" anchor excludes every death the filing's own masthead counts, against a scope statement that forbids exactly that exclusion. And the instrument's evidence exists even though the record never gathered it: Holford et al., JAMA 2014;311(2):164–171 — 8.0 million premature deaths averted, 157 million life-years, 1964–2012 |
Sensitivity (equal weights; full working in steelman-log.md §S5).
| Reading | Top of board |
|---|---|
| 1 — as published | A3 20 · A5 18 · A10 18 · A11 18/7 (mean 2.571) · A7 17 |
| 2 — Phase 1's three contradicted cells applied | A5 18 · A10 18 · A11 18/7 (2.571) · A3 17 · A7 17 |
| 3 — Phase 2 evidence added, A3 scored on the §1115 route that exists | A3 18 · A5 18 · A10 18 · A11 18/7 (2.571) · A7 17 · A9 16 |
A3's first place exists in exactly one of the three readings — the uncorrected one. Under both readings that incorporate work done since publication, "one clear first move" is not what this board says. What survives all three is the comparator: A11 (status quo drift) is at or above every architecture on per-dimension mean in every reading. The filing reports that in a parenthesis; it is the scorecard's most stable result.
The four weightings are not reproducible from this record
This file names four objective weightings (mortality-first, order-first, liberty-first, equal) and the whitepaper and sources page report that CM "ranks first or tied-first under 3 of 4." No weight vector appears in any committed file, and no file contains four ranked lists — here, in ws10-rescore-log.md, or in ws11-findings.md. Only the equal-weight sums can be reproduced, which is why every reading above is computed at equal weights. Deviation #13 discloses that the weights were set by the synthesizing session's judgment rather than pre-registered; what it does not say is that they were never written down, so a reader cannot check the "3 of 4" claim and a re-scorer cannot reproduce it. The owed blinded re-score should publish its weight vectors.
What this scorecard does not yet do (flagged per M6, not smoothed over)
No independent second scorer has reviewed these cellsCompleted 2026-08-04 (ws10-rescore-log.md);no written red-team passcompleted 2026-08-03 (red-team-log.md).- Dimension weights within each objective-weighting scenario were set by the synthesizing session's judgment, not pre-registered before evidence collection (a genuine M3 deviation — see the deviations log). A fully compliant version of this scorecard would have fixed the weighting scheme in the original protocol, before Phase 0.
- Several cells rest on workstream findings that were themselves flagged as moderate-confidence or single-source (e.g., A5's harm-reduction mortality-contribution cell rests on §3's finding that naloxone's causal contribution is asserted-not-isolated).