GUBMENTPlain talk · policy frontier
Filings / Artificial Intelligence / Sources & data
Sources & Data · Series GBMT-SE1 · Filed 2026-08-04

The research record behind Special Edition No. SE1

Every verified anchor, every workstream's finding, and every logged deviation from the protocol — the full record behind "We're regulating a quantity nobody measures."

Record 0

Protocol & scope

United States; the governance of artificial intelligence — frontier-model development and deployment, labor-market effects, the compute/energy build-out, liability for AI harms, and federal-vs-state regulatory architecture. Military and autonomous-weapons AI is excluded (different decision-makers, classified evidence base). AGI-timeline forecasting is excluded as a primary object: this filing scores policy architectures under uncertainty about capabilities; it does not adjudicate capability forecasts. Imports the gubment method (M1–M9) in full.

Phase 0 verdict: GO

Why "Special Edition." The numbered filings each take a policy area with a settled public argument and a measurable binding constraint. AI is different in kind: the object of regulation changes faster than a research cycle, it cuts across every other filing's domain, and the honest confidence ceiling on forward-looking claims is lower. A Special Edition carries the same evidence discipline with an elevated uncertainty register — it stamps UNKNOWABLE AS OF FILING where the numbered series would demand resolution, and uses that stamp without embarrassment.

Record 1

Anchor table — priors, stated before evidence, verified after

Every anchor was written down as an unverified guess before research began, so it could be broken. One fired the filing's kill condition. Two broke. Six of twelve were starred for the Phase 0 gate; the rest were filled or left blank during full execution rather than guessed.

#Anchor (unverified prior)Verified value & delta
1 ★Census puts measured US business AI use in the mid-single-digit to low-teens percent rangeBroken — stale and too low. The recurring series runs 17.3% (Nov 2025) to 21.5% (Jul 2026), already above "low-teens" at its start and still climbing; state range 9.2% (WV) to 28.8% (DC). Single-root caveat: no independent second national instrument was found.
2 ★Private frontier-AI capex exceeds total federal AI spending by >10×Broken as a single ratio — the denominator choice is the finding. ~8–11× if federal spending means contract ceilings ($91.8B, 98.9% DoD); ~90–250× if it means obligated spend ($3–8B). Both readings are defensible and give opposite pictures of state involvement.
3 ★No national instrument measures AI deployment in consequential decisionsCONFIRMED — the kill condition fires as the filing's headline. Hiring, lending, and housing: none. Healthcare: partial and weakening — and thinner than this filing first recorded. Federal certification exposes predictive-algorithm source attributes to a limited set of identified users inside the deploying organisation, not on any public registry, and never deployment volume; HTI-5 proposes removing those paragraphs outright.
4 ★The adjudicated harm record is dominated by discrimination and fraud, not novel frontier harmsDirectionally confirmed with a recast: novel/frontier harms have ≈0 adjudicated US cases. But discrimination leads by settlement count and fraud/CSAM by incident volume — incompatible units, reported separately rather than netted into one ranking.
5Measured labor displacement is large in exposed niches, near-zero in aggregateSupported. No aggregate BLS/Fed break; real niche declines in freelance writing and translation (two independent teams). One outlier flagged: 13–16% relative employment decline for early-career workers in exposed occupations (ADP payroll microdata).
6 ★Announced data-center capacity exceeds utility-contracted capacity by a large multipleConfirmed directionally at 5–25× depending on utility and measure; no clean national ratio exists. The best primary aggregator (a national lab's large-load review) returned 403 to automated fetch — logged, not smoothed over.
7Federal AI posture reversed substantially between 2024 and 2026Confirmed and sharpened: EO 14110 rescinded Jan 2025; two preemption attempts failed (Senate 99–1; NDAA rider dropped); a Dec 2025 order created a DOJ litigation task force that intervened against Colorado's law.
8 ★EU AI Act in force with phased application; cost/relocation evidence thin and industry-sourcedConfirmed and strengthened. The high-risk tier slipped after industry lobbying; the one independent source (the Commission's own research centre) doesn't attribute relocation to the Act at all. Sharpened 2026-08-10 (Phase 2), on the Official Journal text: the vehicle is Regulation (EU) 2026/1744 of 8 July 2026 (OJ 24 July 2026), and it sets two dates, not one — Annex III high-risk to 2 December 2027, Annex I high-risk to 2 August 2028. It amends none of Articles 51–55, 88–94 or 101 and leaves Article 113's point (b) untouched, so the general-purpose-AI obligations and the Commission's evaluation, access and fining powers over frontier labs are unaffected by the delay.
9Illinois BIPA moved firm behavior more than any federal AI guidanceConfirmed on documented settlements ($650M Meta plus a company-wide face-recognition shutdown; Google $100M; TikTok $92M; Snap $35M; Clearview's nationwide sales ban). The comparative claim rests on the absence of a federal counterexample, not a formal study. A 2024 amendment and a Seventh Circuit ruling weaken per-scan damages going forward.
10Frontier labs fund a material share of the AI-safety advocacy/evaluation ecosystemDocumented as a disclosure column, not a capture claim: an industry-funded safety fund grants to researchers evaluating its funders' models; the largest US AI-policy centre was built on philanthropic and industry money; one lab runs a $200M external research fund on AI's labor effects.
11Expert catastrophic-risk estimates span 3+ orders of magnitude with survey selection problemsSupported, with the instruments named: a 2023 researcher survey gives median 5% / mean 16.2% on extinction; a 2026 Delphi study uses a broader harm definition and is not comparable. Named individual estimates run from near-zero to near-certainty. No standardized elicitation method exists.
12No deployed mechanism could verifiably detect a lab violating its safety commitmentsCONFIRMED. The federal evaluation body was renamed and narrowed with voluntary participation; independent evaluators negotiate access per engagement and never hold standing rights to weights, training data, or compute logs. No case was found of an outside body catching a lab's published safety claim as false.
Record 2

Workstream findings

§2 · BaselineThe one federal AI instrument measures the wrong thing

Committed pipeline for the Census Business Trends and Outlook Survey AI supplement — 18 biweekly waves plus state, sector, and firm-size cross-sections, pulled from static files after confirming the series is absent from the Census API entirely. It measures general business adoption, not deployment in any consequential decision, and so cannot close the gap anchor 3 identifies.

baseline/, pull_btos_ai.py, baseline-v0.md

§3 · Present harms & liabilityTwo obstacles, not one — and the case law isn't accumulating

Discovery/opacity is real and contested, with split 2026 rulings going both ways. But product liability's fixed-design-at-sale premise is an independent doctrinal fault line discovery reform wouldn't fix, and the leading generative-AI harm case settled before any ruling. Verdict recorded as indeterminate-leaning-supportive rather than forced to a clean answer the ~20–25-case record can't support.

ws03-adjudication-criteria.md, ws03-findings.md

§4 · LaborMost jobs forecasts have two parents

The circulating "% of jobs exposed" figures trace to two methodological lineages; the WEF employer survey is the one clearly independent instrument. Consequence applied per pre-registered rule: labor architectures are scored against measured displacement only, forecasts quarantined to a labeled speculative register.

ws04-adjudication-criteria.md, ws04-findings.md

§5 · Compute & energyThe largest measurable public cost isn't in any federal budget

Ratepayer cost-shift documented from regulatory proceedings, not advocacy estimates: $50–60B certified in Georgia (with a simultaneous rate freeze as countermeasure), $29.4B in data-center-attributed PJM capacity charges across four auctions. Large-load tariffs exist in three states but are prospective-only and untested against a stranded-load scenario.

ws05-adjudication-criteria.md, ws05-findings.md

§6 · Frontier riskConfirmed trust-based in the US — and no verdict on the risk itself

H6.1 confirmed for the United States: no US mechanism grants standing verification access. Scope corrected 2026-08-10 (Phase 2) — the unqualified form is refuted by EU AI Act Arts. 91–93 and 101; see Record 7. On the catastrophic-risk question the workstream deliberately issues no verdict, publishing instead a structural map (auditable vs. unverifiable commitment subsets) and a four-item ledger of what evidence would change the assessment.

ws06-adjudication-criteria.md, ws06-findings.md

§7 · FederalismBlocking, not substituting — with a correction to our own first pass

Two failed federal preemption attempts, then litigation. The deepening pass corrected the reconnaissance framing: DOJ's complaint presses Equal Protection claims only, narrower than the private plaintiff's own First Amendment, Commerce Clause and vagueness theories — but it pleads SB24-205's duties inseverable and asks the court to invalidate the whole statute, so the theory is narrow and the target is not. Read at primary tier (ECF 17) in the 2026-08-10 verification pass. California, Texas, and Utah show zero federal enforcement action.

ws07-findings.md

§8 · Civil societyGridlock is architecture-dependent, not universal

H8.1 partially refuted: state-law primacy and disclosure mandates each draw three of four camps; frontier-lab licensure is isolated to the safety camp, with civil-rights and labor holding no located position — harder to organize from than opposition. Funding relationships between labs and the evaluation ecosystem recorded as a COI column for the rest of the filing.

ws08-adjudication-criteria.md, ws08-findings.md

§9 · PrecedentsEverything we'd copy was built for things that hold still

FDA, NEPA, nuclear regulation, and China's algorithm registry, each with disanalogies stated as plainly as analogies. Nuclear supplies the steelman §6 owed — real continuous, independently verified oversight is achievable — inseparably bound to the cost ratchet that helped end American nuclear construction. A pattern surfaced across three domains independently: the fixed-artifact assumption breaks the same way each time.

ws09-adjudication-criteria.md, ws09-findings.md

§10 · ScorecardState-law primacy and disclosure co-lead

Nine architectures × five anchored objectives × six weightings, every cell cited, scored as-implemented rather than as-idealized. After red team and blind re-score: state-law primacy and disclosure co-lead; frontier licensure shares the floor with the voluntary status quo; the comprehensive statute is an all-neutral mid-pack row — an artifact of self-certifying proposals and thin implemented evidence, stated as such rather than as a verdict on the underlying goals. Phase 2 (2026-08-10) contested four cells and showed three of those results are one-cell results — see Record 7 and the sensitivity table in ws10-scorecard.md. Note the structural fact behind them: an architecture with no real implemented instance cannot exceed 3 on any axis under this rubric, so its ceiling is the mid-pack and it can only move down.

ws10-scorecard.md, ws10-red-team-log.md, ws10-rescore-log.md

§11 · SequencingThree of four seed claims survive; one needs amending

Protecting state experimentation, building verifiability infrastructure before licensure, and tying labor instruments to measured triggers all hold. Liability-first needs a correction: pair it with a doctrinal design-defect update, because deploying discovery reform alone runs into the settle-before-precedent pattern §3 documented.

ws11-sequencing.md
Record 3

Deviations log

Nineteen entries. The full table lives in the repository; the ones that changed what the filing can claim:

#DeviationEffect
1A delegated researcher flagged and excluded a circulating "EU AI Act enforcement precedent" source that named no case and cited nothingTreated as a positive control that root-tracing works — and a standing caution that this domain's search results contain fabricated-looking material that would pass a casual citation check
2Anchor 2 resolved as a methodological trap rather than a single ratioBoth readings of "federal AI spending" reported with their denominators stated, rather than picking the one that flatters either framing
6§3's H3.1 criteria were written after its evidence (Phase 0's starred-anchor pass), not beforeDisclosed in the criteria document itself; verdict checked against both a strict and a loose reading to bound post-hoc fitting, and recorded as split rather than resolved to whichever reading was convenient
8
amended #18
The DOJ complaint, federal court dockets, and primary agency documents under the December 2025 order all returned 403 to automated fetch — true at filing, superseded 2026-08-10DOJ's legal theory rested on convergent secondary reporting and said so. Superseded: the 2026-08-10 pass downloaded the whole X. AI LLC v. Weiser docket (ECF 1, 12, 17, 18, 22, 24) from RECAP and EO 14365 from the Federal Register, and read them at primary tier — correcting two of the claims this entry hedged. What remains unfetched is the national lab's large-load review and the Commerce CAISI announcement, which stay the top retrieval targets for a next increment
14The scorecard's draft rankings table did not match its own raw scores in any of six columnsCaught by the red team, then independently hand-recomputed a second time by the primary session before acceptance — the corrected table is verified twice, not accepted from either source alone
15→16#15 flagged that no fully independent re-score had run; #16 closed it (2026-08-06)Blind second scorer corrected eleven cells; comparators separated; scorecard v3 is the published matrix
Record 4

Red team

Eleven attacks against the §10 draft. Nine corrected, one defended with a stated decision rule, one resolved by adding a missing citation rather than changing a score. None dismissed without a documented reason.

Master critique — the six-weighting rankings table matched none of its own raw scores

Corrected. Every column was wrong: wrong leaders, wrong values, invented ties, missed ties. Table regenerated and hand-verified independently before acceptance. The corrected result sharpened rather than reversed the draft's finding — the broken arithmetic had understated how far ahead state-law primacy runs.

Two cells scored the wrong kind of evidence

Corrected. Targeted harm-class statutes had been credited for the size of the problem they would target rather than any demonstrated effect — a proposed instrument scored as if implemented. Separately, the do-nothing comparator quietly held a better harm score than the voluntary status quo its own basis text called identical. Both corrected downward.

The same "self-reported, unverifiable" finding was scored differently in two cells without explanation

Resolved by citation, not re-score. The distinguishing fact — mandatory obligations in force versus voluntary opt-in participation — was real but unstated. Added; the split is defensible once stated and was indefensible while silent.

The feasibility gate applied no stated rule for turning coalition positions into open/closed calls

Flagged, not overridden. The gate labels remain, but the interpretive rule behind them ("mixed" is more organizable than "silent") is now written down so a reader can apply a different rule and reach a different answer.

The steelman for licensure was cited only to be discounted

Corrected. Nuclear's continuous-verification model now appears as an explicit sentence: a non-self-certifying licensure design could plausibly score far higher, and the gap being measured is in what has been proposed, not in what is possible.

Record 5

Independent blind re-score

A second scorer, procedurally blinded to the scorecard, red-team log, and sequencing file, re-scored all forty-five cells from the §3–§9 evidence base. Reconciliation moved eleven cells across four failure families: above-neutral design labels, below-neutral mechanism reasoning, one band-definition violation (frontier licensure's catastrophic-risk cell scored below its own band text), and missed evidence (including the voluntary status quo's documented-false-assurance catastrophic-risk score of 1). Consequence: state-law primacy and disclosure/transparency co-lead; the voluntary status quo now scores worse than doing nothing; frontier licensure shares the floor with the voluntary status quo; axis D's flat-3 missing-labor-architecture finding independently reproduced. Full dispositions: ws10-rescore-log.md (scorecard v3).

Record 6

Phase 1 fact-check — against primary law, by a session that did not write this filing

Run 2026-08-10 under the verification protocol. Every factual assertion on both public pages and in the entire research record was extracted and numbered before any checking began, and each numbered claim carries a verdict. The filing’s own citations were treated as claims to be re-derived, not as verification.

Claims extractedVerdicts recordedUnaddressed
3803800

Ten corrections, ten overstatements, two stale disclosures. The corrections are applied to the whitepaper, to this page, and to the record files behind them — not to a footnote.

CorrectedWasIs
NCMEC AI-linked reports, 20251.5M+400,000+ (NCMEC CyberTipline 2025: “more than 400,000 reports of exploitation with a GAI nexus”). Overstated ~3.75×, on a headline tile and in the abstract
Healthcare algorithm disclosurevendors must “publish” which algorithms exist, on a public registry45 CFR 170.315(b)(11)(v)(A)(1) requires access by “a limited set of identified users”. No public registry exists. This makes the filing’s kill-condition finding stronger
Hospital coverageover 96%96%, from 2021 survey data (ASTP/ONC, 90 FR 60979)
DOJ’s Colorado complaint“targeting one drafting choice in one statute”Narrow theory confirmed (Equal Protection only), but ECF 17 ¶53 pleads SB24-205’s duties inseverable and the prayer asks the court to declare the whole statute invalid
The private plaintiff’s theories“Commerce Clause or preemption theories the private plaintiff is arguing”xAI pleads First Amendment, dormant Commerce Clause and vagueness — no preemption count. Neither party pleads preemption
Frontier licensure’s floorshares the floor “including the catastrophic-risk-weighted one”Ties on four of six weightings. On the catastrophic-risk weighting the voluntary status quo is alone at the bottom (2.05 vs 2.55)
Workstream count11-workstream protocol10 (§2–§11)
IC3 fraud category“AI-enabled fraud losses”IC3 defines the category as “AI Related: Information reported contains a reference to artificial intelligence,” a descriptor “used by IC3 for tracking purposes only.” The $893M total is exact
The safety-index quotationframeworks “lack quantitative thresholds…”; commitments an unreliable proxySource says frameworks “sometimes lack”, and scopes the unreliable-proxy conclusion to Google DeepMind, OpenAI and xAI, not every lab
Ratepayer-vs-federal comparison“rivals or exceeds… under either way of counting”$79–89B combined far exceeds the obligated reading ($3–8B) but approaches rather than exceeds the contract-ceiling reading ($91.8B)

Confirmed at primary tier and unchanged. The 99–1 Senate vote (roll call 00363, 1 Jul 2025, S.Amdt. 2814 to H.R. 1); the CFPB’s 67 withdrawn guidance documents including both algorithm circulars (90 FR 20084); EO 14365 (90 FR 58499); the Colorado enforcement suspension (ECF 24, barring the Attorney General generally, not merely as to xAI); the Workday bias-testing privilege ruling (Mobley ECF 340, 29 May 2026); IC3’s $893,346,472 across 22,364 complaints; the comptroller’s nine-of-twelve misrouted 311 calls; every published Census BTOS figure, reproduced exactly from the committed pipeline; and all thirty-six values of the six-weighting rankings table, recomputed independently for a third time.

What could not be reached. 102 claims are recorded UNVERIFIABLE — this pass could not get a tier deeper than the filing did. The largest concentration is §5’s utility figures: the Georgia PSC certification order, the PJM market monitor’s report and the D.C. bill attribution were all unreachable, and they carry the filing’s fourth headline. The §4 labor literature, the §8 funding map and the §9 precedent cost figures are likewise still at the tier the filing left them. Naming that is the point of the exercise, not a footnote to it.

A disclosure that has gone stale in the filing’s favour. The honesty box records that the DOJ complaint and the federal docket system “refused automated access.” That was true when written. It is no longer: the entire X. AI LLC v. Weiser docket and its filings download freely from RECAP, and were read at primary tier in this pass. Two of the three claims that disclosure was hedging turned out to need correction. Corrected 2026-08-10 (second pass): an independent audit found the stale wording still standing on the whitepaper's honesty box and in deviation #8 above, in the same breath as the newly quoted ECF 17 text. Both now carry a narrowed disclosure naming only what is still actually unobtained — the national lab's data-center large-load review and the Commerce Department's CAISI announcement, neither of which was reached. Logged as deviation #18.

Full ledger, every numbered claim with its verdict and the sources fetched: ai/research/verification-log.md.

Record 7

Phase 2 steelman — the case against this filing’s conclusion, built with equal effort

Run 2026-08-10 in a session separate from the fact-check, per the protocol’s independence rule. The target was not chosen freely: this filing’s own protocol document names the direction it leans and the steelman it therefore owes — “the steelman owed in the unfashionable direction is the verifiable-commitment subset… built with equal effort.” Against ten workstreams and four public parts, what the record actually paid was one paragraph and one bracketed sentence inside a scorecard cell. Verdict: the steelman partially survives, and wins on the scorecard.

Withdrawn or correctedWasIs
Standing access to a frontier model“no mechanism grants an outside party standing access to a lab’s weights, training data, or compute logs” — no country attachedWithdrawn as written; re-scoped to the United States. EU AI Act Art. 92 lets the AI Office evaluate a general-purpose model, lets the Commission “appoint independent experts to carry out evaluations on its behalf,” and lets it request access “through APIs or further appropriate technical means and tools, including source code”; Art. 101(1)(d) prices refusal at 3% of worldwide turnover or €15M. Art. 91 + Annex XI reach training-data provenance and “the computational resources used to train the model (e.g. number of floating point operations).” Applicable 2 August 2026 — two days before this filing published. The Digital Omnibus (Reg. (EU) 2026/1744) amends none of it. Not established: any exercise of the power, and Art. 92(3) names source code, not weights
New York’s RAISE Acta pending proposal that “nominally requires third-party audits” with uncertain durabilityEnacted law, but not yet operative. The Phase 2 pass first read Ch. 699 of the Laws of 2025 (signed 2025-12-19) as the statute; that chapter was repealed and replaced before this filing’s date by A9449 / S8828, signed 2026-03-27 as Ch. 96 of the Laws of 2026, which struck General Business Law art. 44-B §§1420–1425 and substituted a new §§1420–1429 modelled on California’s SB 53. On the current law: the effective date is 1 January 2027, so RAISE was not in force when this filing published on 2026-08-04; penalties are $1M first violation / $3M per subsequent (§1427(1)); and there is no regulator access right at all — §1421(5) instead lets the developer redact and requires it to “describe the character and justification of such redaction,” retaining the unredacted material itself for five years. Reports go to an office within the Department of Financial Services, not the Attorney General. What survives on both versions, checked in full: the word “audit” does not appear anywhere in art. 44-B (nor does “auditor” or “independent”), and the 72-hour incident-reporting duty stands (§1422(3)(a)), with a new 24-hour duty for incidents posing an imminent risk of death or serious injury. That makes this filing’s American safety finding stronger, not weaker — the verdict is unchanged and only the supporting figures moved. This is a correction to this pass’s own correction, logged as deviation #20
Four scorecard cellsscored, uncontestedContested, both readings published. 1C (EU evaluation power), 4E (unmeasured incumbency claim the blind re-score corrected in the identical cell one row up), 6A (credit for the size of the gap addressed, on evidence Phase 1 weakened), 8C (band-1 basis after the quotation was corrected to its true scope)

What the steelman failed to take, said as plainly as what it took. The kill condition stands. The American half of the safety finding stands and is reinforced — California’s SB 53 and New York’s RAISE Act were both read in full at primary tier, and neither grants access to a model, its weights, its training data or its compute logs, nor requires an outside evaluation. “Voluntary commitments score worse than doing nothing” survives its own worst case: at the corrected reading of the cell it rests on, the voluntary status quo still finishes below the do-nothing comparator on all six weightings. And the nuclear existence proof now rests on codified text (10 CFR 50.70’s duty to permit inspection and to house a full-time federal inspector rent-free; 10 CFR 50.72’s immediate-notification duty) rather than on this filing’s own summary of itself.

The scorecard sensitivity, and why it is published rather than applied. Three of the four scorecard claims on the whitepaper are one-cell results: correcting 4E alone ends “frontier licensure shares the floor on four of six weightings” (it then shares it on none); correcting 6A alone ends the co-lead (state-law primacy leads alone on all six); moving 1C on the EU evidence takes the comprehensive federal statute from “mid-pack, all-neutral” to co-leading four weightings and leading the catastrophic-risk weighting outright. A steelman pass is not a third blind scorer, so nothing was re-scored on its own authority — the board stands and its fragility is on the record. The arithmetic is a committed script that reproduces all fifty-four published values, and fails loudly if it cannot. A structurally blinded re-score of those four cells is owed.

Also still owed, and named so nobody assumes otherwise: whether the EU has ever used Article 91 or 92; the implementing act governing how its independent experts are appointed; the Illinois statute text; and the three §5 ratepayer headlines, which both verification phases left at secondary tier.

Full record: ai/research/steelman-log.md; ranking script ai/research/ws10/rank.py; primary texts under method/sources/.

THE RECEIPTS · 10-workstream protocol · Phase 0 gate with a GO verdict and a fired kill condition · pre-registered adjudication criteria per workstream · anchor table preserving broken priors · deviations log (19 entries) · committed pipeline: Census BTOS AI supplement, 18 waves, four cross-sections · red team with 11 logged attacks and dispositions · independent blind re-score correcting eleven cells · Phase 1 fact-check (380 claims) · Phase 2 steelman with a committed ranking script · all public in the repository.