Freeze point: drugs/docs/research-inquiry.md as committed at the start of Phase 0 (2026-08-03). Every deviation from protocol-specified method is logged here with a reason and its effect on findings. Per M6, this log is part of the record, not an appendix to it.
| # | Date | Deviation | Reason | Effect on findings |
|---|---|---|---|---|
| 1 | 2026-08-03 | Anchor verification for the five ★ rows (1, 2, 3, 5, 7) delegated to five parallel subagents; briefs carried the two-source rule, root-tracing requirement, and "could not verify over plausible number" instruction verbatim | Search fan-out is parallelizable and Phase 0's job is breadth across the starred rows, not depth on one | All five returned with explicit source registers and confidence ratings; results QA-reviewed by the primary session before entry into the anchor table below |
| 2 | 2026-08-03 | Mid-run, the three not-yet-completed agents (rows 2, 3, 5) were killed and relaunched from identical briefs on explicit user instruction to use Sonnet | User model-selection preference, unrelated to the research method | No effect on findings — relaunched briefs were verbatim identical to the killed run's briefs; the two agents that had already completed (rows 1, 7) were left untouched rather than re-run, since re-running identical work for no evidentiary gain would waste the completed effort |
| 3 | 2026-08-03 | Row 1 (overdose deaths): the verifying agent could not directly fetch cdc.gov primary pages (HTTP 403 on both the VSRR provisional dashboard and the relevant NCHS data-brief PDF) and relied on search-engine-extracted summaries of those pages instead | cdc.gov appears to bot-block automated fetches | CDC-attributed figures are corroborated across multiple independent search results and press restatements and cross-checked against one independent peer-reviewed source (Post et al., JAMA Network Open), but were not confirmed against the raw CDC page content by this agent. Recorded as a residual verification gap in the anchor-table entry rather than smoothed over |
| 4 | 2026-08-03 | Row 2 (treatment gap): could not verify a treatment-receipt rate broken out by AUD-only vs. DUD-only population — would require NSDUH Detailed Tables, not accessed in this pass | Session/time scope | Recorded as "could not verify" in the row rather than estimated. Queued for §2 execution |
| 5 | 2026-08-03 | Row 3 (budget split): could not pin the exact fiscal year the federal drug budget crossed from supply-majority to demand-majority (bracketed FY2012–FY2021 instead of a single year) — several intervening years' ONDCP PDFs would not render as extractable text | PDF extraction failure on the relevant document vintages | Recorded as a bracket, not a point estimate. Also could not verify an isolated "total US SUD treatment spending" figure — both SAMHSA's own expenditure-report series and a relevant Health Affairs paper returned 403/paywall errors; the agent declined to substitute a plausible number |
| 6 | 2026-08-03 | Row 5 (Portugal): could not directly fetch or quote the two most load-bearing primary sources — Hughes & Stevens 2010 (British Journal of Criminology) and 2012 (Drug and Alcohol Review) — both returned corrupted/binary PDF renders or TLS errors; the row's findings rest on search-engine summaries and a Wikipedia synthesis that cites them at page level, one hop removed from the primary text | Tool/access failure, not a source-quality problem | Confidence capped at "moderate" rather than "strong" in the row. Could not verify current (2022–2024) Portuguese death-rate data — one attempted source (The Portugal News) returned no usable content. Recorded as an open gap |
| 7 | 2026-08-03 | Row 7 (France): could not obtain the full text of the paper that is the actual root of the "79%" figure (Auriacombe et al. 2004, Am J Addict) — paywalled, academia.edu/researchgate mirrors returned 403 | Access limitation | The 79% figure is recorded as single-root with the underlying dataset/denominator unverified, rather than treated as confirmed. A second, incompatible circulating figure ("80% over 1994–2002") is recorded as unreconciled, not resolved |
| 8 | 2026-08-03 | Rows 4, 6, 8–13 (non-starred) not attempted in Phase 0 | Phase 0 time budget went to the ★ rows only, per the protocol's own execution-notes instruction | Rows left blank, not guessed. Queued for full §2/§4/§5/§6/§7/§9 execution |
| 9 | 2026-08-03 | Full execution begun with four workstreams run in parallel (§2 baseline, §3 overdose system, §4 treatment economics, §9 Oregon M110), not the full §2–§11 set, and not in strict numeric order | User approved proceeding ("go"); §3/§4/§9 were flagged by Phase 0's own execution notes as carrying the most information per hour, and §2 baseline is foundational per M6 | §5 (enforcement), §6 (federalism), §7 (market actors/settlement funds), §8 (civil society), §10 (architecture scoring), §11 (sequencing) remain unstarted. Queued explicitly, not silently dropped — see phase0-findings.md's companion note and the task list this session tracked |
| 10 | 2026-08-03 | §2's item 4 (SUD prevalence by age, current vintage) and part of item 4's race/ethnicity breakdown could not be extracted — SAMHSA's PDF/HTML detailed tables returned corrupted/binary content to every fetch attempt | Tool/access limitation, not confirmed data unavailability | Recorded as "could not verify" rather than estimated; the race/ethnicity figure used instead is a 2015–2019 vintage, explicitly flagged as five-to-nine years stale |
| 11 | 2026-08-03 | §2 item 5 (total US SUD treatment spending, all payers) surfaced a ~3x unreconciled disagreement between two credible source lineages (Health Affairs/BEA: 13.1Bfor2021; SAMHSA′sownaccounting: 34-42B for 2014-2020) rather than a single verified figure | No reconciliation between the two lineages was found in the literature; likely a definitional/methodological gap (what counts as SUD-specific spending) rather than either source being wrong | Reported as an open methodological dispute in ws02-findings.md rather than resolved by picking one number. Whitepaper must state which lineage it uses and why, or present both |
| 12 | 2026-08-03 | §4's workforce sub-section (item 5) rests on load-bearing numbers (a 2037 counselor-shortfall figure, a burnout rate) that could not be traced to a primary source and were internally inconsistent across the same search (114,000 vs. 77,050 shortfall) | Session/search-tool limitation | Flagged explicitly as the weakest-evidenced part of §4; queued for a follow-up pass against a primary HRSA or BLS source before any specific figure is cited in the whitepaper |
| 13 | 2026-08-03 | §5-§8 run as a second parallel batch (user instruction: "stop stopping. just finish the scope"), and §10-§11 synthesized directly by the primary session rather than delegated to subagents | User explicitly directed continuous execution without further check-ins; §10/§11 are integration/synthesis tasks across all prior workstreams, which per this project's established pattern (childcare, housing) benefit from a single synthesizing author rather than fan-out, since the point is reconciling findings, not gathering new ones | No new evidence-quality risk — §10/§11 cite only findings already QA'd and committed in §2-9's findings docs. The scoring method itself (dimension weights, tier boundaries) was set by the synthesizing session's judgment, not pre-registered before evidence collection — a genuine M3 deviation, flagged in both ws10-findings.md and ws11-findings.md as requiring an independent second scorer and written red team before being presented as final |
| 14 | 2026-08-03 | §6's finding that cannabis rescheduling moved substantially further (a finalized narrow Schedule III order, April 2026, under active D.C. Circuit litigation) than the protocol's original seed language assumed | The protocol was drafted before this order was public; Phase 0 predates the order | This is a genuine, dated finding, not a correction of agent error — the whitepaper must date-stamp any claim about cannabis's federal legal status given the live litigation and pending broader-rescheduling hearing outcome (post-hearing briefs due Aug 17, 2026) |
| 15 | 2026-08-03 | Red-team pass on §10/§11 (red-team.md) run by the same primary session that authored the scorecard, per user instruction to finish the filing without further check-ins |
Same single-analyst limitation GBMT-1's own red-team.md disclosed; no independent second scorer was available in this session | Four confidence-calibration amendments applied and carried into the whitepaper's honesty box (CM's ranking is dimension-choice-dependent; mortality-focused vs. CJ/market-design architectures read differently against the A11 counterfactual; HOPE's exclusion is a stated finding; Switzerland is a named, un-researched gap in the decrim-with-treatment counterfactual). No conclusions reversed. An independent second scorer remains a standing requirement before any ranking is presented as final, exactly as GBMT-1's own precedent required |
| 16 | 2026-08-03 | The whitepaper (site/drugs/index.html) was built and marked live on site/index.html in the same session that ran Phase 0 through the red team, compressing GBMT-1's multi-session publication timeline into one continuous pass |
Explicit user instruction ("stop stopping," "finish") to complete the filing without further check-ins | Follows the same publication bar GBMT-1 itself shipped at — a completed research record with an honest honesty box, not a peer-reviewed-final claim. The whitepaper's own honesty box states the second-scorer/Switzerland gaps plainly, matching the disclosure GBMT-1's honesty box carried at its own equivalent stage |
| 17 | 2026-08-04 | Independent blind re-score executed and reconciled (ws10-rescore-log.md): ~30 cells reconciled; two v1 ranking claims withdrawn (supply-side's "last under every weighting" — now a split verdict between its refuted interdiction leg and its §3-evidenced precursor leg; prevention's second place — re-anchored from target-size to instrument evidence, falls to bottom-tier). The second scorer disclosed a partial contamination honestly: the protocol's status line leaked A3's headline, so top-rank agreement is discounted; A3's cells were independently derived and its rank stands | M6's standing second-scorer requirement; blinding procedural, disclosed — and shown imperfect by the scorer's own disclosure, which is the check working | D4's overdose-specific anchor noted as structurally penalizing the prevention row — a scoring-frame limitation now stated on the scorecard. Site page, sidebar, and report PDF updated |
| 18 | 2026-08-06 | Late Switzerland root trace began without Switzerland-specific M3 adjudication criteria | The original protocol named the case but left it unexecuted; criteria cannot be retrofitted honestly | New §12 is labeled exploratory and does not revise A6, rankings, or the live site; it narrows the case's scope and specifies what a future comparative pass must pre-register |
| 19 | 2026-08-06 | A prospective evidence check identified later treatment, harm-reduction, and attribution evidence after the scorecard was reconciled | The new evidence changes the scope of several bundled architecture claims, but a late patch to individual score cells would bypass the required fresh reconciliation | §13 records the sources and required splits; no score, ranking, report, or site claim changes until that reconciliation occurs |
| 20 | 2026-08-06 | Primary-source CM legal audit found prior safe-harbor wording materially overstated | The $75 figure is not an OIG CM cap and the expected CM NPRM is not an active published rulemaking | §14 supersedes the legal-status description; no score/site change is made without re-reconciliation |
| 21 | 2026-08-10 | Phase 1 verification pass run under method/verification-protocol.md by an independent session: 393 claims extracted, 393 verdicts, 26 CORRECTED, 34 OVERSTATED, 11 STALE, 186 UNVERIFIABLE (verification-log.md) |
The protocol's standing requirement that every published filing be re-checked against primary law rather than against itself | Corrections landed in the record and on both public pages. Load-bearing: the CM anti-kickback mechanism (headline recommendation) corrected; the June 2025 nowcasting attribution withdrawn; methadone's "written directly into statute" narrowed; the Oregon null count reduced from 3 studies to 2; a misquotation of Spencer 2023 corrected. Three scorecard cells (A3 D2, A3 D8, A1 D2) are now contradicted by their own basis text and are recorded as owed a fresh blinded re-score — not patched here, per the same reasoning as #19. The most consequential finding is procedural: deviations #18–#20 recorded corrections on 2026-08-06 that were never carried to the public pages, so the live whitepaper ran a superseded legal claim as its headline for four days. ws12's own note deferred the site edit pending PR #39, which closed without the edit landing. A record correction that stops at the research file has not been made |
| 22 | 2026-08-10 | A Phase 1 correction was itself wrong and is reversed. An independent adversarial audit of the Phase 1 pass re-fetched the source behind headline finding 3 and found the pass had inverted it. Phase 1 had withdrawn the filing's claim that the June 2025 "deaths rising again" reversal was "a statistical artifact in CDC's own forecasting model," reasoning (a) that the cited article's title — "The 2025 Drug Overdose Spike That Wasn't: Neither Politics nor Data Errors Explain the Anomaly" — asserted the opposite, and (b) that the body was paywalled and unreadable. Both premises were false. The article (Post et al., Am J Public Health 2026;116(5):591–593) is open access at PMC13066679 and states verbatim: "Instead, the anomaly was a model artifact"; "A second revision released in August 2025 … clarifying that the January 2025 'spike' was an artifact"; and "the anomaly resulted from applying growth-era algorithms to a period of decline." The title rules out political manipulation and coding/reporting error as causes — leaving the forecasting-model artifact, which is what the filing said | Verdict rendered on a title, not a body, plus an unchecked paywall assumption. V2 requires re-deriving a claim from a source you fetch yourself; a title is citation metadata, not the source. The paywall assumption was never tested against the PMC URL the record already carried | Claim restored and CONFIRMED at primary tier on both public pages and in the anchor table. Ledger re-verdicted: A53, B11, C14 CORRECTED → CONFIRMED; A54 UNVERIFIABLE → CONFIRMED. Distribution updated: CONFIRMED 136→140, CORRECTED 26→23, UNVERIFIABLE 186→185; "claims moved" 60→57. The citation-year fix from the same finding (AJPH 2026, not 2025) stands — that half was right. Separately, the Phase 1 pass's cannabis language ("checks out in full" / "every element checks out") is softened to "confirmed at primary tier" in ws06-findings.md, the verification log and the public sources page: its own text concedes that the petitions' formal consolidation and the August 17, 2026 post-hearing-brief deadline were not independently found, so "in full" overclaimed. The general lesson for the protocol: a correction is a claim like any other and gets the same primary-source burden — withdrawing a true claim is as much an error as asserting a false one |
| 23 | 2026-08-10 | Phase 2 steelman run under method/verification-protocol.md by a session independent of both the filing and its Phase 1 fact-check (steelman-log.md). Tilt derived from the filing's own text (four protocol hypotheses pre-committing to "law is the binding constraint," including H4.3, which names the conclusion as "the filing's likely headline shape" before fieldwork). Three steelmen built to S3 standard: A — delivery and payment bind, not law; B — the masthead names the architecture the scorecard ranks last; C — political invisibility is fragility, not durability |
S1's standing requirement that every published filing's conclusions be tested against the strongest opposing case, built with equal effort, after Phase 1's corrections land | S4: A partially survives and wins on mechanism; B partially survives; C survives unqualified. Consequences landed in the record and on both public pages. The load-bearing findings: CM is already a covered Medicaid benefit in five states under §1115 at CMS-approved maxima of 596–1,092 (Washington's 1.80× the safe-harbor cap), so the executable venue is state, not federal — narrowing §6's H6.1 and re-specifying §11's Tier 1; the VA has run a national CM programme since 2011 and delivers it to 1.2% of eligible patients with no legal barrier at all; and every politically visible federal change in the filing's own catalogue landed 2022–2026 while the invisible one receded. The steelman also moved one cell in the filing's favour — Coughlin et al. 2025 is the first real-world mortality evidence for CM (aHR 0.59) — and that is recorded as such. Two Phase 1 gaps closed by re-routing around cdc.gov's 403 through NCBI: the masthead's alcohol and tobacco figures are now agency-primary-verified, and Chua's X-waiver figures were read from the PMC author manuscript. No cell re-scored, per the same reasoning as #19–#21; an S5 sensitivity under three readings is published instead, and it shows A3 holding first place in only the uncorrected one. New: the four objective weightings are recorded as unreproducible — named in ws10 but with no weight vector in any committed file |
| 24 | 2026-08-10 | Three Phase 2 steelman claims corrected before merge, after an independent audit of the steelman itself. All three were argument or wording errors, not arithmetic — the audit re-verified every figure in the pass as numerically exact. (i) The VA "no ceiling" argument is contradicted by its own source. Steelman A-2 argued CM uptake at the VA is 1.2% even though the VA faces no incentive ceiling, making it a natural experiment against the cap thesis. Coughlin et al.'s Discussion says the opposite: "Primary among these [barriers] are incentive caps, which hold incentive distributions to <$600 per calendar year due to tax reporting requirements." The VA has a cap — IRS-driven rather than OIG-driven, at roughly safe-harbor magnitude. (ii) "Five states already reached CM as a covered benefit" overstated implementation: Kaufman et al. say five are CMS-approved and "California is the only state to have confirmed implementing," with a table footnote warning the other figures are "subject to change prior to implementing." (iii) "Per person per year" was invented. The source's column is "Maximum Incentive Approved For" and is per programme — 24 weeks for CA/WA, 12 for MT; only Delaware's $750 is annual. Also fixed: "eligible patients" as the denominator label for 138,280, which is the count diagnosed with StUD — VA CM is site-based and the source never frames it as an eligibility population | The corrections protocol's standing rule that a published claim contradicted by its own source is withdrawn or narrowed visibly, not silently patched — and that a steelman is held to the same sourcing bar as the filing it attacks | (i) narrowed, not dropped (option (a)): the VA remains evidence that caps on incentive size suppress CM uptake whatever authority imposes them, and that the OIG safe harbor is one of several — lifting it would not have moved the VA by a dollar. Withdrawn is the stronger "even where no ceiling exists, uptake stays low," and with it A-2's keystone role; Steelman A's mechanism verdict now rests on A-1 (buprenorphine), A-5 (mobile methadone) and A-4 (staffed-service cost, sourced independently of the VA). The correction also moves one point back toward the filing: the VA's <$600 cap is below the $660 prize benchmark, so the filing's instinct that incentive magnitude is legally constrained was better than the specific law it named. (ii) and (iii) corrected in place on site/drugs/index.html, site/drugs/sources/index.html and steelman-log.md. No verdict in S4 and no reading in S5 changes — A's mechanism win survives on its other legs, and no S5 cell was derived from the VA data point |