GUBMENTPlain talk · policy frontier
Filings / Disability / Sources / §9 Candidate Architecture Scorecard (G
GBMT-12 · Research record · No. 12

§9 Candidate Architecture Scorecard (GBMT-12)

disability/research/ws09-scorecard.md
This is a working research document from the disability filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 12, not a summary of it.

Date: 2026-08-11. Pass 2 — structurally blinded re-score reconciled (ws09-rescore-log.md; raw output ws09-blind-scores-2026-08-11.md). Prior: pass 1 + M6 red-team (ws09-red-team-log.md). Batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx (claude-opus-5 via scripts/batch-rescore.py). Method: Per M6/M10 — anchored ordinal scales authored first (outcome/measured-property anchors, never instrument labels); every cell cites a workstream finding or is held at 3; rankings under the three protocol weightings with explicit weight vectors; rank-stability reported. Symmetric floor (housing/elder-care/media correction): below-neutral cells require cited evidence of an adverse or undermining result — absence of research is 3, not 1 or 2. Above-neutral cells likewise require cited evidence of the measured property — mechanism inference alone is 3. KC1 fires: every O1 cell that would need a point fraud rate inherits a band (written as the cell value plus the band note in the honesty box).

Objectives (protocol §2): O1 measurement integrity · O2 adjudication timeliness and accuracy · O3 poverty and insurance protection · O4 work without a cliff · O5 administrability.


Anchored scales (1–5)

Authored before scoring. Anchors describe outcomes / measured properties, not architecture names (M6).

O1 — Measurement integrity

Whether the architecture’s public case and its operating statistics measure the same object speakers claim.

Score Anchor
1 Direct evidence the architecture depends on or amplifies a mislabeled statistic (e.g. treats stewardship improper-payment % as intentional fraud).
2 Relies on contested constructs with a documented mislabel or construct-mix risk, without correcting it.
3 No project evidence of improvement or harm to measurement honesty; or the cell inherits the KC1 band (cannot claim a point intentional-fraud rate).
4 Analogous / partial evidence that the integrity object matches a separable administrative series (CDR cessations, wage/SGA reporting) rather than fraud theater.
5 Direct evidence the architecture installs or uses a clean intentional-fraud (or correctly labeled) series at the claimed magnitude.

O2 — Adjudication timeliness and accuracy

Time to decision / hearing and residual error or unexplained variance after case mix.

Score Anchor
1 Direct evidence of lengthening waits or worsening unexplained variance / remand churn.
2 Documented mechanism that consumes scarce DDS/ALJ capacity or raises appeal volume without a measured timeliness/accuracy gain.
3 Unevidenced or mixed in this record.
4 Analogous / partial evidence of wait reduction or residual-variance narrowing under quality/capacity interventions.
5 Direct measured clearance of SSA’s own wait standard, or a large residual-variance reduction, under the instrument.

O3 — Poverty and insurance protection

Poverty / deep poverty and health-coverage continuity for beneficiaries and denied applicants.

Score Anchor
1 Direct evidence of coverage or cash loss (or poverty increase) attributable to the instrument.
2 Documented partial adverse protection effect (eligibility cuts, long unprotected waits) without offsetting coverage gains.
3 Unevidenced or mixed.
4 Analogous / partial evidence of poverty reduction or coverage continuity under the instrument class.
5 Direct causal / strong quasi-experimental evidence of protection gains.

O4 — Work without a cliff

Ability to attempt work without a one-step cash + Medicaid/Medicare exit; population-scale earnings/employment effects.

Score Anchor
1 Direct evidence the instrument steepens the cliff or reduces work attempts via benefit design.
2 Primary evaluation finds null or ≪2pp population employment/earnings effects for the lever the architecture claims, or take-up stays negligible while the cliff remains.
3 Unevidenced or mixed (including bundles whose arms conflict).
4 Analogous / partial evidence that coverage-separation or cliff simplification moves earnings/work more than information-only tools.
5 Direct causal evidence of large population employment gains from cliff redesign.

O5 — Administrability

Whether SSA / DDSs can execute the instrument at current or plausible capacity.

Score Anchor
1 Direct evidence of implementation collapse or sustained non-execution.
2 Requires capacity or financing preconditions the record shows are missing or non-portable, with a cautionary failed/turbulent precedent.
3 Unevidenced or mixed at current capacity.
4 Analogous evidence of executable rollout inside existing agency tools / state options.
5 Direct evidence of successful execution at material scale under current or recently demonstrated capacity.

Scores (pass 2 — reconciled)

Pass-2 cell moves vs post-red-team marked ‡ (ws09-rescore-log.md). Prior red-team moves (†) retained in history only.

# Architecture O1 O2 O3 O4 O5 Basis
1 Adjudication capacity surge (DDS/ALJ hiring, digitization; clear backlog before tightening) 3 4 3 3 4 O2=4: H2 supported — hearing APT 342 days (FY2024) vs 270-day goal (FY2023 450); pending hearings ~262k (−19% YoY); initial pending ~1.18M, initial APT 231 (ws02). Not 5: no RCT of a named surge; cause of FY2024 hearing improvement not isolated. O3=3: wait-clearance-as-hardship remains mechanism inference (red-team Attack 2; blind matched). O5=4: staffing/pending are the standing management response; APT/pending trajectory shows the lever moves. O1/O4=3.
2 ALJ consistency / quality controls (cut unexplained variance without blanket crackdown) 4‡ 4 3 3 4 O1=4‡ (was 3): GAO-18-37 residual dispersion after case-mix + ~5 pp narrowing under QA/training is a separable administrative series; row refuses raw-allowance mislabel (ws02) — construct alignment parallel to #6/#8. O2=4 / O5=4: same QA/outlier-review evidence. O3/O4=3.
3 Cliff redesign (1 − for2 offsets, simplified earnings reporting, automatic 1619(b)/Medicare extensions) 3 3 3 2 4‡ O4=2 against BOND: Stage 1/2 null mean earnings; Stage 1 benefits +143 * */yr, Stage2 * *+450–500/yr; EWIC null (ws04, ws05). O5=4‡ (was 3): BOND operated a national offset for years inside SSA tools; 1619(b)/extended Medicare already exist — anchor 4. O3 kept 3 (judgment split): benefits-due rise is largely SGA windfall, not poverty/coverage evidence (rescore log). O1/O2=3.
4 Medicaid buy-in expansion tied to disability work attempts 3 3 4 4 4 Blind matched cell-for-cell. O4=4: WA MBI matched evaluation + Mathematica MBI series (ws04, ws07). Not BOND-cleared cash-offset. O3=4: coverage continuity. O5=4: 47 states already offer a buy-in (KFF). O1/O2=3.
5 Ticket / VR redesign or replacement 3 3 3 2 3 Blind matched. O4=2: ITT STW/employment null or undetectable (≪2pp); service enrollment +0.1–0.4pp; assignment ~1–5% (ws05). O1/O2/O3/O5=3.
6 Integrity focused on CDRs and wage reporting (not fraud theater) 4 3 2‡ 3 4 O1=4: AFR stewardship causes; CDR cessation ≠ original-award fraud (ws03). O3=2‡ (was 3): cessations at scale (~39k disabled-worker FY2019; ~136k DDS pre-COVID) are documented cash/coverage exits — anchor 2 protection effect; not a fraud finding (ws03, ws06). O5=4: CDR volumes ≫ CDI judicial actions **~74–77**/year. O2/O4=3.
7 Partial disability / graduated benefits (VA-like or Nordic-inspired) 3 3‡ 3 3 2 O2=3‡ (was 2): UK ESA/PIP caution is assigned to #9, not #7; demotion instruction is O5/O4 only (ws07) — elder-care floor. O5=2: NL WIA / Nordic stack preconditions missing in US DI. O4=3: no US DI earnings evaluation of a graduated schedule. O1/O3=3.
8 Childhood SSI separate track 4 3 3 3 4 O1=4: adult IP-as-fraud and Ticket/SGA frames mis-travel (ws06, ws08). O3=3: caseload drivers ≠ measured protection gains from the architecture (red-team Attack 4; blind matched). O5=4 kept (judgment split vs blind 3): PRWORA already executed a child-specific standard (ws07). O2/O4=3.
9 Definition tightening / listings reform (“get tougher”) 2 2 2 3 2 Blind matched all five. O1=2: advocacy rides IP-as-fraud citogenesis without correcting it (red-team Attack 5). O2=2 / O5=2: capacity-blind tightening under ~1.18M pending initials; UK reassessment caution (ws02, ws07). O3=2: eligibility-cut / hardship precedent (PRWORA-adjacent). O4=3.
10 Do-nothing comparator 3‡ 3‡ 3 2 3 O1=3‡ / O2=3‡ (were 2): citogenesis is downstream of correctly labeled AFR stewardship; hearing APT improved while initials worsened — mixed → 3 (ws02, ws08). O4=2: cliffs unchanged; 1619(b) 2.6%; Ticket ≪2pp; BOND null. O3/O5=3.

Explicit weight vectors

Protocol §3 readers, fixed before ranking:

Weighting Reader (w_{O1}) (w_{O2}) (w_{O3}) (w_{O4}) (w_{O5})
Applicant in the backlog Waiting on DDS / ALJ; life on hold 0.10 0.40 0.25 0.05 0.20
Beneficiary who might work Wants earnings without losing Medicaid / one-way exit 0.15 0.10 0.25 0.40 0.10
Integrity hawk / budget staffer Rolls bloated or agency cannot police them 0.30 0.20 0.10 0.10 0.30

Score for architecture (a) under weighting (w) is (\sum_i w_i \cdot s_{a,i}). Unweighted mean (=\frac{1}{5}\sum_i s_{a,i}) reported for transparency only — it is not a protocol reader.


Rankings (pass 2)

Rank Applicant in backlog Score Beneficiary who might work Score Integrity hawk / budget Score
1 #2 ALJ QC 3.70 #4 Medicaid buy-in expansion 3.75 #2 ALJ QC 3.80
2 #1 Capacity surge 3.60 #2 ALJ QC 3.35 #8 Childhood separate track 3.60
3 #4 Medicaid buy-in expansion 3.50 #8 Childhood track 3.25 #1 Capacity = #4 Buy-in = #6 CDR/wage 3.50
4 #8 Childhood track 3.30 #1 Capacity surge 3.20
5 #3 Cliff redesign 3.15 #6 CDR/wage 3.00
6 #5 Ticket = #10 Do-nothing 2.95 #7 Partial disability 2.90 #3 Cliff redesign 3.20
7 #3 Cliff 2.70 #5 Ticket = #10 Do-nothing 2.90
8 #7 Partial disability 2.80 #5 Ticket = #10 Do-nothing 2.60 #7 Partial disability 2.70
9
10 #9 Definition tightening 2.05 #9 Definition tightening 2.40 #9 Definition tightening 2.10

Unweighted means (not a reader): #2 = #4 = 3.60 · #1 = #8 = 3.40 · #6 = 3.20 · #3 = 3.00 · #5 = #7 = #10 = 2.80 · #9 = 2.20.

Tops under each weighting (headline)

Weighting Top architecture
Applicant in the backlog #2 ALJ consistency / QC (pass-1/red-team #1=#2 co-lead broken)
Beneficiary who might work #4 Medicaid buy-in expansion (unchanged)
Integrity hawk / budget #2 ALJ consistency / QC (pass-1/red-team #6=#8 co-lead broken; #8 second as construct artifact)

Demoted fraud-first result

A prosecution-centric / fraud-theater integrity program is not among the ten scored rows because the protocol already replaced it with #6. Ghost-row scores (red-team Attack 8), provisional: O1=1, O2=3, O3=3, O4=3, O5=2 → applicant 2.60 / beneficiary 2.60 / integrity hawk 2.10. Strictly below #6 under every weighting. Grounding: KC1; CDI judicial actions ~74–77/year vs CDR cessations ~39k–100k+ (ws03).

#9 (definition tightening) is last under all three weightings on the pass-2 board — because #10’s unevidenced below-neutral cells returned to 3, not because #9 was cut further (blind matched #9’s entire row). This is not a quiet reassertion of the red-team’s withdrawn overclaim; it is a pass-2 arithmetic consequence of the evidence floor on do-nothing (rescore log).


Rank-stability

Rank-stable top set (top four under every weighting): #1 (capacity surge), #2 (ALJ QC), #4 (Medicaid buy-in), #8 (childhood separate track). Post-red-team’s #1/#4/#6 set does not survive: #6’s O3 demotion exits the applicant top four; #2’s O1 lift enters every reader’s top band.

Stable low set: #9 last under all three on this board; #5 and #10 tied in the lower half under beneficiary and mid-low elsewhere; #3 lifted on O5 but still bound by BOND on O4; #7 mid-low once O2 is neutralized and O5=2 remains.

What is not stable: a single “reform disability” winner. The three protocol readers still pick different first-place instruments only in the weak sense that beneficiary stays on #4 while backlog and integrity both land on #2 after pass 2 — a tighter top than pass 1, still not a monopoly.

Sensitivity (one cell): if #3’s automatic-1619(b) arm were scored like #4 on O4 (lift O4 2→4), #3’s beneficiary score rises 2.70→3.50 — still behind #4 (3.75). The BOND null remains load-bearing for any cash-offset-led cliff story.


Honesty box

  1. KC1 — O1 is a band, not a point. FY2023 SSI improper payments 10.62% and OASDI ~0.30% are stewardship (nonmedical) estimates, not intentional-fraud rates (ws03, Phase 0). Every O1 cell inherits this band. Row #6’s O1=4 means construct alignment with CDR/wage-reporting series; #8’s O1=4 means construct separation for child SSI; #2’s O1=4 means case-mix-honest residual-variance measurement — none is a measured fraud reduction.

  2. BOND null constraint; three different work levers. Cliff redesign (#3) scored against BOND: null mean earnings, small average benefit increases, Stage 1 net social cost (ws04). Cliff redesign ≠ Ticket redesign ≠ Medicaid buy-in. Ticket (#5) remains ITT ≪2pp (ws05). Buy-in (#4) attacks the coverage cliff — it is not “the cliff redesign that cleared the BOND bar.”

  3. Awards ≠ stock (H8). Mental disorders are 12.7% of 2023 disabled-worker awards and ≈28.6% of disabled-worker stock (ws08). No architecture is scored as if psychiatric awards were a one-third fraud scandal.

  4. Twice-checked, not final (M10). Pass 1 authored scales and cells in one session; red-team amended three cells; this pass structurally blinded the board (42/50 match; 6 moves; 2 judgment splits). Per M10, one blind re-score is “twice checked,” not settled forever.

  5. Symmetric floor applied (both directions). Pass-2 moves that enforce it: #7 O2 and #10 O1/O2 returned to 3; #6 O3 moved to 2 on cited cessations; #3 O3 not lifted on benefits-due windfall. Judgment splits logged in the rescore log.

  6. #8’s integrity-hawk second place remains a construct artifact. Refusing adult-fraud imports is an O1/O5 win; it is not the adult DI Trust Fund fix. Adult integrity / budget readers should read #2 (measurement-honest adjudication QC) and #6 (CDR/wage series) as the adult-facing integrity instruments — #6 no longer co-leads after the O3 demotion.

← All Disability research documents Sources digest Read the whitepaper