Date: 2026-08-11. Method: per M6 and the drugs/housing precedents — written attacks on the committed pass-1 ws09-scorecard.md, each checked against the workstream record, each given a verdict. Amendments applied in the same revision of the scorecard; this log records what changed and why. Blind re-score still owed (not run in this pass).
Eight attacks: 3 SUSTAINED (cell moves), 3 PARTIALLY SUSTAINED (framing/honesty), 2 REJECTED.
Attack log
Attack 1 — Same-session scale + score circularity
Verdict: PARTIALLY SUSTAINED — cannot be fully rebutted from inside (same as GBMT-1 Attack 5 / drugs Attack 1). Anchored scales and cell scores were authored in one session. Rank stability across weightings therefore measures internal consistency under a self-chosen frame, not external truth. Mitigations already present (outcome anchors, not architecture names; every non-neutral cell cites a workstream finding; symmetric floor) reduce but do not eliminate the risk.
Action: Honesty box now names the circularity explicitly and treats the structurally blinded re-score as the standing requirement before ranks are “twice checked.” No cell move.
Attack 2 — #1 capacity surge evidence quality (O2/O3/O5 = 4)
Verdict: SUSTAINED on O3; REJECTED on O2 and O5.
- O2=4 holds. Hearing APT 450→342 with pending −19% YoY while still above the 270-day goal is measured wait improvement under capacity easing (ws02). That meets the O2=4 anchor (analogous / partial evidence of wait reduction). Not 5: no RCT of a named surge — already disclosed.
- O5=4 holds. Staffing/pending are the agency’s standing management response, and the same trajectory shows the lever moves inside existing tools.
- O3=4 does not hold. The cell rested on H7 process structure (“clearing waits is insurance-against-hardship for eventually-allowed claimants”) — mechanism inference, not measured poverty / coverage outcomes under a surge. Above-neutral cells require cited evidence of the measured property (elder-care / housing evidence-floor lesson applied upward). Absence of a poverty series is 3, not a free 4.
Action: #1 O3 4→3. Applicant score 3.85→3.60 (now ties #2); beneficiary 3.45→3.20; integrity 3.60→3.50.
Attack 3 — #4 buy-in vs #3 cliff redesign given the BOND null
Verdict: PARTIALLY SUSTAINED on framing; REJECTED on cell moves.
The board correctly separates instruments: #3’s O4=2 is the BOND cash-offset null on mean earnings; #4’s O4=4 is coverage-separation evidence (WA matched MBI; Mathematica MBI series) against information-only nulls (ws04). That is not “ignoring BOND” — it is refusing to let a DI cash-offset RCT veto a different cliff (health coverage). Residual risk: readers may still hear “cliff redesign failed, buy-in worked” as the same lever with opposite results.
Action: Honesty box restates that #4 is not BOND-cleared cliff redesign and that its O4=4 is quasi-experimental / descriptive, not RCT-grade. Ticket (#5) remains a third row. No cell move. Sensitivity on auto-1619(b) inside #3 unchanged in direction.
Attack 4 — #8 childhood track leading the integrity-hawk reader: artifact?
Verdict: SUSTAINED — artifact, with a cell move that removes the solo lead.
Two defects stacked: (1) O3=4 cited caseload drivers (poverty, school evidence, Medicaid) — baseline description of who is on child SSI, not evidence that a separate-track architecture improves poverty/protection outcomes; (2) the integrity-hawk weighting loads O1 at 0.30, so a construct-correction row (adult frames mis-travel) outranks adult CDR integrity on measurement honesty alone. Honesty box §6 already warned; the solo #8 lead still over-claimed.
Action: #8 O3 4→3. Integrity-hawk score 3.70→3.60 — now ties #6 (CDR/wage integrity) rather than leading alone. Applicant score drops 3.55→3.30 (exits the old top-four set). Adult-roll readers should treat #6 as the integrity co-lead; #8 remains a construct win, not a Trust Fund fix.
Attack 5 — #9 “last under all weightings” and the elder-care absence/adverse lesson
Verdict: SUSTAINED on O1; headline “last under all three” dies.
O2=2 / O3=2 / O5=2 cite adverse mechanisms (capacity consumption, eligibility hardship, UK reassessment caution) — those survive the symmetric floor. O1=1 did not. Anchor 1 requires the architecture to depend on or amplify a mislabeled statistic as its operating object. The record shows get-tougher advocacy rides IP-as-fraud citogenesis (ws08); listings reform as an instrument can be pursued without that mislabel. Scoring the campaign’s rhetoric as the instrument’s O1=1 converted discourse into adverse evidence — the elder-care failure mode.
Action: #9 O1 1→2 (contested construct / mislabel risk without correcting it, as the architecture is typically sold on this board). Scores: applicant 1.95→2.05; beneficiary 2.25→2.40; integrity 1.80→2.10. Under beneficiary weighting, #10 do-nothing is now last (2.35); #9 is last only under applicant and integrity hawk. Replace “last under all three” with the narrower true claim.
Attack 6 — KC1 band on O1
Verdict: PARTIALLY SUSTAINED — disclosure, not re-banding every cell. KC1 correctly forbids reading any O1 point as an intentional-fraud rate. Row #6’s O1=4 is construct alignment; #8’s O1=4 is construct separation; neither is a fraud-reduction claim. Point ordinals remain the scorecard’s grammar; the band lives in the honesty box.
Action: Honesty box tightened — every O1 cell inherits the band; no O1 point score may be cited as a fraud %. No cell move.
Attack 7 — Ticket vs cliff conflation
Verdict: REJECTED. Pass-1 already scored #3 against BOND and #5 against Ticket ITT as separate rows; honesty box forbade collapsing them; the auto-1619(b) arm inside #3 is disclosed as usability, not as a BOND clear. Attack 3’s framing amendment is the residual reader risk; no further cell change.
Attack 8 — Fraud-theater omission hides a competing architecture?
Verdict: REJECTED as a hide; PARTIALLY SUSTAINED as opacity of the ghost scores. Protocol candidate list already replaced prosecution-centric fraud theater with #6 (deviation 16). Omitting an 11th row is not a quiet kill if the counterfactual is published. Pass-1 asserted “last or near-last” without showing the arithmetic.
Action: Publish explicit ghost-row scores in the scorecard demotion section: provisional O1=1, O2=3, O3=3, O4=3, O5=2 → applicant 2.60 / beneficiary 2.60 / integrity 2.10 — below #6 under every weighting; on integrity hawk, tied-with-or-below corrected #9 and below do-nothing. No new scored row; no cell move on the ten.
Cell moves (pass-1 → post-red-team)
| Cell | Before | After | Attack |
|---|---|---|---|
| #1 O3 | 4 | 3 | 2 |
| #8 O3 | 4 | 3 | 4 |
| #9 O1 | 1 | 2 | 5 |
Ranking changes (headline)
| Weighting | Pass-1 top | Post-red-team top |
|---|---|---|
| Applicant in the backlog | #1 (3.85) | #1 = #2 (3.60) |
| Beneficiary who might work | #4 (3.75) | #4 (3.75) — unchanged |
| Integrity hawk / budget | #8 (3.70) | #6 = #8 (3.60) |
Rank-stable top set (top four under every weighting): was #1/#4/#8 → now #1 / #4 / #6 (#8 exits applicant top four; #6 enters all three).
#9 last under all three: withdrawn. #9 last under applicant (2.05) and integrity hawk (2.10); under beneficiary, #10 is last (2.35).
What did not change
Workstream findings (H1–H8, KC1–KC4, BOND null, Ticket ≪2pp, CDR ≫ prosecutions, awards ≠ stock) were not attacked as false and are not revised. Tops still hold in the weak sense: capacity instruments lead the backlog reader; Medicaid buy-in alone leads might-work; measurement-honest integrity work (#6/#8) leads the integrity hawk — with #8’s solo lead demoted to a tie and named as a construct artifact. Blind re-score still owed before “twice checked.”