Date: 2026-08-11. Pass 2 — structurally blinded re-score reconciled (ws09-rescore-log.md; raw output ws09-blind-scores-2026-08-11.md). Prior: pass 1 + M6 red-team (ws09-red-team-log.md). Batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx (claude-opus-5 via scripts/batch-rescore.py). Method: Per M6/M10 — anchored ordinal scales authored first (outcome/measured-property anchors, never instrument labels); every cell cites a workstream finding or is held at 3; rankings under the three protocol weightings with explicit weight vectors; rank-stability reported. Symmetric floor (housing/elder-care/media correction): below-neutral cells require cited evidence of an adverse or undermining result — absence of research is 3, not 1 or 2. Above-neutral cells likewise require cited evidence of the measured property — mechanism inference alone is 3. KC1 fires: every O1 cell that would need a point fraud rate inherits a band (written as the cell value plus the band note in the honesty box).
Objectives (protocol §2): O1 measurement integrity · O2 adjudication timeliness and accuracy · O3 poverty and insurance protection · O4 work without a cliff · O5 administrability.
Anchored scales (1–5)
Authored before scoring. Anchors describe outcomes / measured properties, not architecture names (M6).
O1 — Measurement integrity
Whether the architecture’s public case and its operating statistics measure the same object speakers claim.
| Score | Anchor |
|---|---|
| 1 | Direct evidence the architecture depends on or amplifies a mislabeled statistic (e.g. treats stewardship improper-payment % as intentional fraud). |
| 2 | Relies on contested constructs with a documented mislabel or construct-mix risk, without correcting it. |
| 3 | No project evidence of improvement or harm to measurement honesty; or the cell inherits the KC1 band (cannot claim a point intentional-fraud rate). |
| 4 | Analogous / partial evidence that the integrity object matches a separable administrative series (CDR cessations, wage/SGA reporting) rather than fraud theater. |
| 5 | Direct evidence the architecture installs or uses a clean intentional-fraud (or correctly labeled) series at the claimed magnitude. |
O2 — Adjudication timeliness and accuracy
Time to decision / hearing and residual error or unexplained variance after case mix.
| Score | Anchor |
|---|---|
| 1 | Direct evidence of lengthening waits or worsening unexplained variance / remand churn. |
| 2 | Documented mechanism that consumes scarce DDS/ALJ capacity or raises appeal volume without a measured timeliness/accuracy gain. |
| 3 | Unevidenced or mixed in this record. |
| 4 | Analogous / partial evidence of wait reduction or residual-variance narrowing under quality/capacity interventions. |
| 5 | Direct measured clearance of SSA’s own wait standard, or a large residual-variance reduction, under the instrument. |
O3 — Poverty and insurance protection
Poverty / deep poverty and health-coverage continuity for beneficiaries and denied applicants.
| Score | Anchor |
|---|---|
| 1 | Direct evidence of coverage or cash loss (or poverty increase) attributable to the instrument. |
| 2 | Documented partial adverse protection effect (eligibility cuts, long unprotected waits) without offsetting coverage gains. |
| 3 | Unevidenced or mixed. |
| 4 | Analogous / partial evidence of poverty reduction or coverage continuity under the instrument class. |
| 5 | Direct causal / strong quasi-experimental evidence of protection gains. |
O4 — Work without a cliff
Ability to attempt work without a one-step cash + Medicaid/Medicare exit; population-scale earnings/employment effects.
| Score | Anchor |
|---|---|
| 1 | Direct evidence the instrument steepens the cliff or reduces work attempts via benefit design. |
| 2 | Primary evaluation finds null or ≪2pp population employment/earnings effects for the lever the architecture claims, or take-up stays negligible while the cliff remains. |
| 3 | Unevidenced or mixed (including bundles whose arms conflict). |
| 4 | Analogous / partial evidence that coverage-separation or cliff simplification moves earnings/work more than information-only tools. |
| 5 | Direct causal evidence of large population employment gains from cliff redesign. |
O5 — Administrability
Whether SSA / DDSs can execute the instrument at current or plausible capacity.
| Score | Anchor |
|---|---|
| 1 | Direct evidence of implementation collapse or sustained non-execution. |
| 2 | Requires capacity or financing preconditions the record shows are missing or non-portable, with a cautionary failed/turbulent precedent. |
| 3 | Unevidenced or mixed at current capacity. |
| 4 | Analogous evidence of executable rollout inside existing agency tools / state options. |
| 5 | Direct evidence of successful execution at material scale under current or recently demonstrated capacity. |
Scores (pass 2 — reconciled)
Pass-2 cell moves vs post-red-team marked ‡ (ws09-rescore-log.md). Prior red-team moves (†) retained in history only.
| # | Architecture | O1 | O2 | O3 | O4 | O5 | Basis |
|---|---|---|---|---|---|---|---|
| 1 | Adjudication capacity surge (DDS/ALJ hiring, digitization; clear backlog before tightening) | 3 | 4 | 3 | 3 | 4 | O2=4: H2 supported — hearing APT 342 days (FY2024) vs 270-day goal (FY2023 450); pending hearings ~262k (−19% YoY); initial pending ~1.18M, initial APT 231 (ws02). Not 5: no RCT of a named surge; cause of FY2024 hearing improvement not isolated. O3=3: wait-clearance-as-hardship remains mechanism inference (red-team Attack 2; blind matched). O5=4: staffing/pending are the standing management response; APT/pending trajectory shows the lever moves. O1/O4=3. |
| 2 | ALJ consistency / quality controls (cut unexplained variance without blanket crackdown) | 4‡ | 4 | 3 | 3 | 4 | O1=4‡ (was 3): GAO-18-37 residual dispersion after case-mix + ~5 pp narrowing under QA/training is a separable administrative series; row refuses raw-allowance mislabel (ws02) — construct alignment parallel to #6/#8. O2=4 / O5=4: same QA/outlier-review evidence. O3/O4=3. |
| 3 | Cliff redesign (1 − for−2 offsets, simplified earnings reporting, automatic 1619(b)/Medicare extensions) | 3 | 3 | 3 | 2 | 4‡ | O4=2 against BOND: Stage 1/2 null mean earnings; Stage 1 benefits +143 * */yr, Stage2 * *+450–500/yr; EWIC null (ws04, ws05). O5=4‡ (was 3): BOND operated a national offset for years inside SSA tools; 1619(b)/extended Medicare already exist — anchor 4. O3 kept 3 (judgment split): benefits-due rise is largely SGA windfall, not poverty/coverage evidence (rescore log). O1/O2=3. |
| 4 | Medicaid buy-in expansion tied to disability work attempts | 3 | 3 | 4 | 4 | 4 | Blind matched cell-for-cell. O4=4: WA MBI matched evaluation + Mathematica MBI series (ws04, ws07). Not BOND-cleared cash-offset. O3=4: coverage continuity. O5=4: 47 states already offer a buy-in (KFF). O1/O2=3. |
| 5 | Ticket / VR redesign or replacement | 3 | 3 | 3 | 2 | 3 | Blind matched. O4=2: ITT STW/employment null or undetectable (≪2pp); service enrollment +0.1–0.4pp; assignment ~1–5% (ws05). O1/O2/O3/O5=3. |
| 6 | Integrity focused on CDRs and wage reporting (not fraud theater) | 4 | 3 | 2‡ | 3 | 4 | O1=4: AFR stewardship causes; CDR cessation ≠ original-award fraud (ws03). O3=2‡ (was 3): cessations at scale (~39k disabled-worker FY2019; ~136k DDS pre-COVID) are documented cash/coverage exits — anchor 2 protection effect; not a fraud finding (ws03, ws06). O5=4: CDR volumes ≫ CDI judicial actions **~74–77**/year. O2/O4=3. |
| 7 | Partial disability / graduated benefits (VA-like or Nordic-inspired) | 3 | 3‡ | 3 | 3 | 2 | O2=3‡ (was 2): UK ESA/PIP caution is assigned to #9, not #7; demotion instruction is O5/O4 only (ws07) — elder-care floor. O5=2: NL WIA / Nordic stack preconditions missing in US DI. O4=3: no US DI earnings evaluation of a graduated schedule. O1/O3=3. |
| 8 | Childhood SSI separate track | 4 | 3 | 3 | 3 | 4 | O1=4: adult IP-as-fraud and Ticket/SGA frames mis-travel (ws06, ws08). O3=3: caseload drivers ≠ measured protection gains from the architecture (red-team Attack 4; blind matched). O5=4 kept (judgment split vs blind 3): PRWORA already executed a child-specific standard (ws07). O2/O4=3. |
| 9 | Definition tightening / listings reform (“get tougher”) | 2 | 2 | 2 | 3 | 2 | Blind matched all five. O1=2: advocacy rides IP-as-fraud citogenesis without correcting it (red-team Attack 5). O2=2 / O5=2: capacity-blind tightening under ~1.18M pending initials; UK reassessment caution (ws02, ws07). O3=2: eligibility-cut / hardship precedent (PRWORA-adjacent). O4=3. |
| 10 | Do-nothing comparator | 3‡ | 3‡ | 3 | 2 | 3 | O1=3‡ / O2=3‡ (were 2): citogenesis is downstream of correctly labeled AFR stewardship; hearing APT improved while initials worsened — mixed → 3 (ws02, ws08). O4=2: cliffs unchanged; 1619(b) 2.6%; Ticket ≪2pp; BOND null. O3/O5=3. |
Explicit weight vectors
Protocol §3 readers, fixed before ranking:
| Weighting | Reader | (w_{O1}) | (w_{O2}) | (w_{O3}) | (w_{O4}) | (w_{O5}) |
|---|---|---|---|---|---|---|
| Applicant in the backlog | Waiting on DDS / ALJ; life on hold | 0.10 | 0.40 | 0.25 | 0.05 | 0.20 |
| Beneficiary who might work | Wants earnings without losing Medicaid / one-way exit | 0.15 | 0.10 | 0.25 | 0.40 | 0.10 |
| Integrity hawk / budget staffer | Rolls bloated or agency cannot police them | 0.30 | 0.20 | 0.10 | 0.10 | 0.30 |
Score for architecture (a) under weighting (w) is (\sum_i w_i \cdot s_{a,i}). Unweighted mean (=\frac{1}{5}\sum_i s_{a,i}) reported for transparency only — it is not a protocol reader.
Rankings (pass 2)
| Rank | Applicant in backlog | Score | Beneficiary who might work | Score | Integrity hawk / budget | Score |
|---|---|---|---|---|---|---|
| 1 | #2 ALJ QC | 3.70 | #4 Medicaid buy-in expansion | 3.75 | #2 ALJ QC | 3.80 |
| 2 | #1 Capacity surge | 3.60 | #2 ALJ QC | 3.35 | #8 Childhood separate track | 3.60 |
| 3 | #4 Medicaid buy-in expansion | 3.50 | #8 Childhood track | 3.25 | #1 Capacity = #4 Buy-in = #6 CDR/wage | 3.50 |
| 4 | #8 Childhood track | 3.30 | #1 Capacity surge | 3.20 | — | — |
| 5 | #3 Cliff redesign | 3.15 | #6 CDR/wage | 3.00 | — | — |
| 6 | #5 Ticket = #10 Do-nothing | 2.95 | #7 Partial disability | 2.90 | #3 Cliff redesign | 3.20 |
| 7 | — | — | #3 Cliff | 2.70 | #5 Ticket = #10 Do-nothing | 2.90 |
| 8 | #7 Partial disability | 2.80 | #5 Ticket = #10 Do-nothing | 2.60 | #7 Partial disability | 2.70 |
| 9 | — | — | — | — | — | — |
| 10 | #9 Definition tightening | 2.05 | #9 Definition tightening | 2.40 | #9 Definition tightening | 2.10 |
Unweighted means (not a reader): #2 = #4 = 3.60 · #1 = #8 = 3.40 · #6 = 3.20 · #3 = 3.00 · #5 = #7 = #10 = 2.80 · #9 = 2.20.
Tops under each weighting (headline)
| Weighting | Top architecture |
|---|---|
| Applicant in the backlog | #2 ALJ consistency / QC (pass-1/red-team #1=#2 co-lead broken) |
| Beneficiary who might work | #4 Medicaid buy-in expansion (unchanged) |
| Integrity hawk / budget | #2 ALJ consistency / QC (pass-1/red-team #6=#8 co-lead broken; #8 second as construct artifact) |
Demoted fraud-first result
A prosecution-centric / fraud-theater integrity program is not among the ten scored rows because the protocol already replaced it with #6. Ghost-row scores (red-team Attack 8), provisional: O1=1, O2=3, O3=3, O4=3, O5=2 → applicant 2.60 / beneficiary 2.60 / integrity hawk 2.10. Strictly below #6 under every weighting. Grounding: KC1; CDI judicial actions ~74–77/year vs CDR cessations ~39k–100k+ (ws03).
#9 (definition tightening) is last under all three weightings on the pass-2 board — because #10’s unevidenced below-neutral cells returned to 3, not because #9 was cut further (blind matched #9’s entire row). This is not a quiet reassertion of the red-team’s withdrawn overclaim; it is a pass-2 arithmetic consequence of the evidence floor on do-nothing (rescore log).
Rank-stability
Rank-stable top set (top four under every weighting): #1 (capacity surge), #2 (ALJ QC), #4 (Medicaid buy-in), #8 (childhood separate track). Post-red-team’s #1/#4/#6 set does not survive: #6’s O3 demotion exits the applicant top four; #2’s O1 lift enters every reader’s top band.
Stable low set: #9 last under all three on this board; #5 and #10 tied in the lower half under beneficiary and mid-low elsewhere; #3 lifted on O5 but still bound by BOND on O4; #7 mid-low once O2 is neutralized and O5=2 remains.
What is not stable: a single “reform disability” winner. The three protocol readers still pick different first-place instruments only in the weak sense that beneficiary stays on #4 while backlog and integrity both land on #2 after pass 2 — a tighter top than pass 1, still not a monopoly.
Sensitivity (one cell): if #3’s automatic-1619(b) arm were scored like #4 on O4 (lift O4 2→4), #3’s beneficiary score rises 2.70→3.50 — still behind #4 (3.75). The BOND null remains load-bearing for any cash-offset-led cliff story.
Honesty box
KC1 — O1 is a band, not a point. FY2023 SSI improper payments 10.62% and OASDI ~0.30% are stewardship (nonmedical) estimates, not intentional-fraud rates (ws03, Phase 0). Every O1 cell inherits this band. Row #6’s O1=4 means construct alignment with CDR/wage-reporting series; #8’s O1=4 means construct separation for child SSI; #2’s O1=4 means case-mix-honest residual-variance measurement — none is a measured fraud reduction.
BOND null constraint; three different work levers. Cliff redesign (#3) scored against BOND: null mean earnings, small average benefit increases, Stage 1 net social cost (ws04). Cliff redesign ≠ Ticket redesign ≠ Medicaid buy-in. Ticket (#5) remains ITT ≪2pp (ws05). Buy-in (#4) attacks the coverage cliff — it is not “the cliff redesign that cleared the BOND bar.”
Awards ≠ stock (H8). Mental disorders are 12.7% of 2023 disabled-worker awards and ≈28.6% of disabled-worker stock (ws08). No architecture is scored as if psychiatric awards were a one-third fraud scandal.
Twice-checked, not final (M10). Pass 1 authored scales and cells in one session; red-team amended three cells; this pass structurally blinded the board (42/50 match; 6 moves; 2 judgment splits). Per M10, one blind re-score is “twice checked,” not settled forever.
Symmetric floor applied (both directions). Pass-2 moves that enforce it: #7 O2 and #10 O1/O2 returned to 3; #6 O3 moved to 2 on cited cessations; #3 O3 not lifted on benefits-due windfall. Judgment splits logged in the rescore log.
#8’s integrity-hawk second place remains a construct artifact. Refusing adult-fraud imports is an O1/O5 win; it is not the adult DI Trust Fund fix. Adult integrity / budget readers should read #2 (measurement-honest adjudication QC) and #6 (CDR/wage series) as the adult-facing integrity instruments — #6 no longer co-leads after the O3 demotion.