Brief: attack the scorecard synthesis's conclusions as hard as the evidence allows, per M6 and the GBMT-1 precedent (childcare/research/red-team.md). Each attack gets a response; where the attack lands, the conclusion is amended, not defended.
Attack 1: The scorecard's dimension weights were set by the same session that scored the cells
Same single-analyst problem GBMT-1's Attack 5 identified. If the scorer chose both the dimensions and the scores, "rank stability across weightings" measures internal consistency, not external truth. Response: cannot be fully rebutted from inside — same as GBMT-1. Mitigations in place (anchored scales defined before scoring, every cell cites a specific workstream finding) reduce but don't eliminate the risk. Standing requirement, already flagged in ws10-findings.md: an independent second scorer must review the cells before this ranking is presented as anything more than a first-pass structure.
Attack 2: A3 (contingency management)'s top ranking may be an artifact of which dimensions were chosen, not a discovery about which architecture is best
The dimension set (legal executability, cost, speed) structurally favors narrow regulatory fixes over ambitious, institution-spanning architectures like A6 (decriminalization) or A9 (the tobacco-playbook transplant). A narrow fix will almost always out-score a broad one on "how fast/cheap is this," regardless of the broad one's actual value if achieved. Response: partially lands. The dimension set was a deliberate choice reflecting this project's cross-filing pattern (GBMT-1's workforce finding, GBMT-2's legal-capacity finding) that regulatory/legal capacity is usually the binding constraint — but that's a design choice carried in from other filings, not a discovery this filing re-derived independently. Amendment: the whitepaper must state explicitly that the objective weightings tested don't include a "scale/ambition of impact" axis. CM's ranking is strongest under a framework that values immediately-executable levers; a reader who weights transformative-scale change over speed-to-execution could reasonably prefer a different architecture. State CM as "the clearest first move," not "the recommendation" — a real narrowing of the claim.
Attack 3: §3 found the 2023-25 overdose decline is supply-side/exposure-driven, largely independent of policy — doesn't that undercut the premise that any architecture here matters much?
If deaths are already falling for reasons no policy lever controls, why build an elaborate architecture scorecard at all? Response: partially lands, and is already flagged in §11 as the kill-condition-adjacent finding — A11 (status quo drift) is a genuinely strong comparator specifically for this reason. But it doesn't undercut the whole exercise: (a) the decline could reverse — the documented June 2025 nowcasting false-alarm episode (§3) already shows the public/media are primed to believe a reversal on thin evidence, meaning policy readiness matters regardless of the trend's current direction; (b) mortality is only one of the outcomes this filing scores against — the CJ-footprint, treatment-access-desert, and market-design problems documented in §4/§5/§7 (rehab-industry fraud, cannabis illicit-market persistence, buprenorphine pricing disparities) exist independent of which theory explains the mortality trend. Amendment: state explicitly in the whitepaper that mortality-focused architectures (A2, A3, A5) should be read against the A11 counterfactual; architectures targeting CJ-footprint or market-design problems don't depend on which overdose-decline theory is correct.
Attack 4: Excluding HOPE-style supervision from the enforcement steelman (A10) may itself be tilted, given M3's warning against under-crediting enforcement
The protocol explicitly requires an equal-effort steelman for enforcement (M3). Did dropping HOPE — the most famous swift-certain-fair model — quietly weaken the very steelman the protocol demanded? Response: does not land. The exclusion is evidence-based, not tilt-driven: §5 found the higher-quality, later evidence (a four-site multi-site replication RCT) showed no advantage over conventional probation, directly contradicting the original single-site result HOPE's popular reputation rests on. Focused deterrence and drug courts — both retained in A10 — have real, replicated positive evidence in the trial record. This is the M3-mandated steelman working as designed: built from the actual evidence, not from an intervention's popular reputation. No amendment; note in the whitepaper that HOPE's exclusion is itself a finding worth stating plainly, not a quiet omission.
Attack 5: Portugal and Oregon are both used as cautionary "treatment-funding-delay" cases for A6 — is this finding unfalsifiable, since any decrim case without fast treatment funding will always look like it failed for that reason?
Without a comparison case where decriminalization was paired with prompt, adequate treatment funding, "fund treatment fast" is an untested prescription dressed as a finding. Response: partially lands, and identifies a real gap. §9 only root-traced Portugal and Oregon in this pass; the protocol's own seed list named Switzerland (heroin-assisted treatment plus the "four pillars" model, often cited as the case that did fund treatment adequately) as a case queued but not yet researched. Amendment: the whitepaper's honesty box must state plainly that A6's "fund treatment fast" prescription currently rests on two cases sharing the same failure mode, and that the positive counterfactual (Switzerland) is an identified, un-researched gap — a priority item for the next research pass, not a settled comparison.
Net effect on §10/§11
No conclusions reversed. Four confidence-calibration amendments applied, all logged above and to be carried into the whitepaper's honesty box: (1) CM's top ranking is dimension-choice-dependent — restate as "clearest first move," not "the recommendation"; (2) mortality-focused architectures must be read against the A11 status-quo counterfactual, while CJ-footprint/market-design architectures don't depend on that trend's attribution; (3) HOPE's exclusion from the enforcement steelman is a stated finding, not an omission; (4) A6's treatment-funding-delay lesson rests on two same-failure-mode cases — Switzerland is a named, real gap, not yet closed. Per M6, the scorecard's core structure (contingency management as the standout first move; supply-side modernization as the weakest performer; decriminalization and enforcement as the two axes of largest rank instability) survives this pass and is strengthened by having its weakest points named rather than defended.