Date: 2026-08-04 Method: Per M6 — a second scorer (separate agent, procedurally blinded: scale definitions + as-implemented architecture list + the ws02–ws10 evidence base only; the scorecard, red-team log, sequencing file, and all site pages forbidden) scored all 55 cells independently. Reconciled cell-by-cell below. Blinding is procedural and disclosed.
Headline: the durability axis was still leaking design labels after the red team explicitly forbade exactly that
GBMT-9's first red team renamed axis B to "demonstrated durability against a hostile political majority, not a design label" — and corrected two cells accordingly. The blind re-score shows the correction didn't go far enough: five more B cells still scored above the track-record floor on design reasoning (a 2024-enacted credit at 3; never-implemented vouchers at 3; a 2023 philanthropy fund at 4). Under the axis's own anchors — 1 = "no demonstrated track record," 3 = "has meaningfully outlasted normal political cycles" — a brand-new instrument cannot score 3. All five corrected.
Two further systematic catches, both familiar shapes:
- Below-neutral-without-evidence violations (the same asymmetric-floor failure the housing re-score exposed the same day): five cells sat at 2 on mechanism reasoning with no project evidence. Corrected to 3.
- The comparators had real evidence the original scored as vibes. The do-nothing and cash rows' concentration and trust cells are backed by verified record findings (the Peterson & Dunaway ownership result, the two-country polarization mechanism, the live consolidation trajectory) — the blind scorer cited them cell-by-cell where the original had treated comparator scores as structural filler.
Material reconciliations (all 55 cells were compared; identical cells omitted)
| Cell | Orig | Blind | Resolved | Disposition |
|---|---|---|---|---|
| 1B | 3 | 1 | 1 | NY/IL/NM credits enacted 2024-26; no track record exists. Blind's basis cited the CA deal's collapse, which the 2026-08-04 moving-targets pass corrected (survived at ~8× reduced scale) — noted, but immaterial to a no-record instrument's floor |
| 1C | 3 | 4 | 4 | Illinois's actual disbursement data (120+ outlets, ~30% to nonprofits, mostly outside Chicago) is practice evidence the original under-credited |
| 1D | 2 | 3 | 3 | Below-neutral-without-evidence violation |
| 2A | 2 | 1 | 1 | The ~90%-to-three-incumbents figure plus Canada's local-outlet devastation is the A-axis anchor-1 case exactly |
| 2B | 2 | 1 | 1 | Meta's walk-away is a completed failed real-world test |
| 2D | 2 | 1 | 1 | An antitrust exemption with verified concentration of proceeds |
| 2E | 2 | 3 | 3 | No trust evidence either way → neutral |
| 3B | 3 | 2 | 2 | Partial real record: Maryland's tax base survived litigation; the earmark has no record at all |
| 4C / 4bC | 4 | 5 | 5 | Distribution to ~1,500 stations including the highest-dependency tribal stations is anchor-5 "demonstrated in practice" |
| 4D / 4bD | 2 | 3 | 3 | Below-neutral violations; no concentration evidence |
| 4bE | 2 | 3 | 3 | De-partisanizing mechanism is plausible, unverified → 3 (4E stays 2 — both scorers independently cited the partisan-fight evidence) |
| 5A | 2 | 3 | 3 | H11.1 is an untested seed hypothesis — below-neutral violation |
| 5B | 3 | 1 | 1 | Never implemented anywhere; no record |
| 5C | 4 | 3 | 3 | "Viewpoint-neutral by construction" is a design label; C wants demonstrated practice |
| 5E | 4 | 3 | 3 | Same design-label leak |
| 6B | 2 | 1 | 2 | Judgment split kept: NJCIC has 8 years of survival and a demonstrated 60% cut in routine budget politics — the middle of a thin record. Logged |
| 6D | 2 | 3 | 3 | Below-neutral violation |
| 7A | 1 | 3 | 3 | The original conflated scale-insufficiency (the real, evidenced ~55× gap) with no capacity added; philanthropy demonstrably adds capacity (INN's 443-outlet sector), just far below need. The 55× constraint stays in the basis text |
| 7B | 4 | 1 | 2 | Both corrected: original's 4 was design reasoning ("voluntary"); blind's 1 ignores philanthropy's long class-level independence record. The instrument (2023 fund + public match) has no record, and its match component inherits the CPB failure precedent → 2 |
| 7C | 3 | 4 | 3 | No Press Forward distributional breakdown exists; sector-composition inference isn't demonstrated allocation. Kept 3; logged |
| 7D | 2 | 3 | 3 | Below-neutral violation |
| 8A | 2 | 3 | 3 | "No desert-targeting mechanism" is reasoning, not evidence |
| 8E | 4 | 3 | 3 | "Plausibly narrows the gap" is the definition of plausibility — the E axis's own hard rule |
| 9C | 1 | 2 | 2 | The countervailing verified growth (+144% digital-native jobs, INN sector) belongs in the cell |
| 10C | 3 | 2 | 2 | Comparator inherits the verified concentration trajectory |
| 10D | 1 | 2 | 2 | One notch above do-nothing on a flagged-unverified offset; kept below neutral on the verified trajectory |
Reconciled matrix
| # | Architecture | A | B | C | D | E |
|---|---|---|---|---|---|---|
| 1 | Payroll tax credit | 3 | 1 | 4 | 3 | 3 |
| 2 | Platform bargaining code | 1 | 1 | 1 | 1 | 3 |
| 3 | Ad tax, earmarked | 3 | 2 | 3 | 3 | 3 |
| 4 | Public media, appropriated | 4 | 1 | 5 | 3 | 2 |
| 4b | Public media, non-appropriated | 4 | 3 | 5 | 3 | 3 |
| 5 | News vouchers/credits | 3 | 1 | 3 | 3 | 3 |
| 6 | State civic-info consortia | 3 | 2 | 4 | 3 | 3 |
| 7 | Philanthropy match | 3 | 2 | 3 | 3 | 3 |
| 8 | Structural antitrust | 3 | 3 | 3 | 3 | 3 |
| 9 | Managed transition | 1 | 5 | 2 | 1 | 1 |
| 10 | Cash | 1 | 5 | 2 | 2 | 2 |
What the reconciliation does to the findings
- The non-appropriated public-media variant (4b) is now the clear most-consistent performer (neutral-weight mean 3.6; nothing else above 3.0) — a strengthened version of the post-red-team finding, now resting on corrected cells rather than idealized ones. Its B=3 remains capped by the Finland counter-example; that cap survived both scorers independently.
- The bargaining code (2) is the record's one affirmatively multi-axis-refuted architecture — 1s on four of five axes, each cell resting on peer-reviewed or platform-measured evidence, not absence of evidence. The old "most stable bottom performer" framing understated this: it is not merely last, it is the only instrument the record actively argues against.
- The comparators are no longer filler. Do-nothing's 1s on supply/competitiveness/trust are verified-trajectory cells (consolidation accelerating; the two-country polarization mechanism operating on the more-trusted local tier). Doing nothing is now the evidenced worst option on three axes, not the assumed one.
- 17 of 55 cells are 3-by-hard-rule — the blind scorer's count, carried forward: most of this scorecard's differentiation is evidence-density, the same meta-finding as housing's same-day reconciliation.
- One blind-scorer input (the CA Google deal "collapse") was already stale against the same-day moving-targets correction — an instance of the moving-target problem inside the re-score itself, disclosed here.
Effects applied
ws11-scorecard.md revised to the reconciled matrix (v3, evidence-coverage noted); whitepaper Part 9 + sidebar updated; report PDF regenerated; deviation appended.