GUBMENTPlain talk · policy frontier
Filings / Disability / Sources / §9 Red-Team Log (GBMT-12)
GBMT-12 · Research record · No. 12

§9 Red-Team Log (GBMT-12)

disability/research/ws09-red-team-log.md
This is a working research document from the disability filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 12, not a summary of it.

Date: 2026-08-11. Method: per M6 and the drugs/housing precedents — written attacks on the committed pass-1 ws09-scorecard.md, each checked against the workstream record, each given a verdict. Amendments applied in the same revision of the scorecard; this log records what changed and why. Blind re-score still owed (not run in this pass).

Eight attacks: 3 SUSTAINED (cell moves), 3 PARTIALLY SUSTAINED (framing/honesty), 2 REJECTED.


Attack log

Attack 1 — Same-session scale + score circularity

Verdict: PARTIALLY SUSTAINED — cannot be fully rebutted from inside (same as GBMT-1 Attack 5 / drugs Attack 1). Anchored scales and cell scores were authored in one session. Rank stability across weightings therefore measures internal consistency under a self-chosen frame, not external truth. Mitigations already present (outcome anchors, not architecture names; every non-neutral cell cites a workstream finding; symmetric floor) reduce but do not eliminate the risk.

Action: Honesty box now names the circularity explicitly and treats the structurally blinded re-score as the standing requirement before ranks are “twice checked.” No cell move.

Attack 2 — #1 capacity surge evidence quality (O2/O3/O5 = 4)

Verdict: SUSTAINED on O3; REJECTED on O2 and O5.

Action: #1 O3 4→3. Applicant score 3.85→3.60 (now ties #2); beneficiary 3.45→3.20; integrity 3.60→3.50.

Attack 3 — #4 buy-in vs #3 cliff redesign given the BOND null

Verdict: PARTIALLY SUSTAINED on framing; REJECTED on cell moves.

The board correctly separates instruments: #3’s O4=2 is the BOND cash-offset null on mean earnings; #4’s O4=4 is coverage-separation evidence (WA matched MBI; Mathematica MBI series) against information-only nulls (ws04). That is not “ignoring BOND” — it is refusing to let a DI cash-offset RCT veto a different cliff (health coverage). Residual risk: readers may still hear “cliff redesign failed, buy-in worked” as the same lever with opposite results.

Action: Honesty box restates that #4 is not BOND-cleared cliff redesign and that its O4=4 is quasi-experimental / descriptive, not RCT-grade. Ticket (#5) remains a third row. No cell move. Sensitivity on auto-1619(b) inside #3 unchanged in direction.

Attack 4 — #8 childhood track leading the integrity-hawk reader: artifact?

Verdict: SUSTAINED — artifact, with a cell move that removes the solo lead.

Two defects stacked: (1) O3=4 cited caseload drivers (poverty, school evidence, Medicaid) — baseline description of who is on child SSI, not evidence that a separate-track architecture improves poverty/protection outcomes; (2) the integrity-hawk weighting loads O1 at 0.30, so a construct-correction row (adult frames mis-travel) outranks adult CDR integrity on measurement honesty alone. Honesty box §6 already warned; the solo #8 lead still over-claimed.

Action: #8 O3 4→3. Integrity-hawk score 3.70→3.60 — now ties #6 (CDR/wage integrity) rather than leading alone. Applicant score drops 3.55→3.30 (exits the old top-four set). Adult-roll readers should treat #6 as the integrity co-lead; #8 remains a construct win, not a Trust Fund fix.

Attack 5 — #9 “last under all weightings” and the elder-care absence/adverse lesson

Verdict: SUSTAINED on O1; headline “last under all three” dies.

O2=2 / O3=2 / O5=2 cite adverse mechanisms (capacity consumption, eligibility hardship, UK reassessment caution) — those survive the symmetric floor. O1=1 did not. Anchor 1 requires the architecture to depend on or amplify a mislabeled statistic as its operating object. The record shows get-tougher advocacy rides IP-as-fraud citogenesis (ws08); listings reform as an instrument can be pursued without that mislabel. Scoring the campaign’s rhetoric as the instrument’s O1=1 converted discourse into adverse evidence — the elder-care failure mode.

Action: #9 O1 1→2 (contested construct / mislabel risk without correcting it, as the architecture is typically sold on this board). Scores: applicant 1.95→2.05; beneficiary 2.25→2.40; integrity 1.80→2.10. Under beneficiary weighting, #10 do-nothing is now last (2.35); #9 is last only under applicant and integrity hawk. Replace “last under all three” with the narrower true claim.

Attack 6 — KC1 band on O1

Verdict: PARTIALLY SUSTAINED — disclosure, not re-banding every cell. KC1 correctly forbids reading any O1 point as an intentional-fraud rate. Row #6’s O1=4 is construct alignment; #8’s O1=4 is construct separation; neither is a fraud-reduction claim. Point ordinals remain the scorecard’s grammar; the band lives in the honesty box.

Action: Honesty box tightened — every O1 cell inherits the band; no O1 point score may be cited as a fraud %. No cell move.

Attack 7 — Ticket vs cliff conflation

Verdict: REJECTED. Pass-1 already scored #3 against BOND and #5 against Ticket ITT as separate rows; honesty box forbade collapsing them; the auto-1619(b) arm inside #3 is disclosed as usability, not as a BOND clear. Attack 3’s framing amendment is the residual reader risk; no further cell change.

Attack 8 — Fraud-theater omission hides a competing architecture?

Verdict: REJECTED as a hide; PARTIALLY SUSTAINED as opacity of the ghost scores. Protocol candidate list already replaced prosecution-centric fraud theater with #6 (deviation 16). Omitting an 11th row is not a quiet kill if the counterfactual is published. Pass-1 asserted “last or near-last” without showing the arithmetic.

Action: Publish explicit ghost-row scores in the scorecard demotion section: provisional O1=1, O2=3, O3=3, O4=3, O5=2 → applicant 2.60 / beneficiary 2.60 / integrity 2.10 — below #6 under every weighting; on integrity hawk, tied-with-or-below corrected #9 and below do-nothing. No new scored row; no cell move on the ten.


Cell moves (pass-1 → post-red-team)

Cell Before After Attack
#1 O3 4 3 2
#8 O3 4 3 4
#9 O1 1 2 5

Ranking changes (headline)

Weighting Pass-1 top Post-red-team top
Applicant in the backlog #1 (3.85) #1 = #2 (3.60)
Beneficiary who might work #4 (3.75) #4 (3.75) — unchanged
Integrity hawk / budget #8 (3.70) #6 = #8 (3.60)

Rank-stable top set (top four under every weighting): was #1/#4/#8 → now #1 / #4 / #6 (#8 exits applicant top four; #6 enters all three).

#9 last under all three: withdrawn. #9 last under applicant (2.05) and integrity hawk (2.10); under beneficiary, #10 is last (2.35).

What did not change

Workstream findings (H1–H8, KC1–KC4, BOND null, Ticket ≪2pp, CDR ≫ prosecutions, awards ≠ stock) were not attacked as false and are not revised. Tops still hold in the weak sense: capacity instruments lead the backlog reader; Medicaid buy-in alone leads might-work; measurement-honest integrity work (#6/#8) leads the integrity hawk — with #8’s solo lead demoted to a tie and named as a construct artifact. Blind re-score still owed before “twice checked.”

← All Disability research documents Sources digest Read the whitepaper