Date: 2026-08-11 Model: claude-opus-5 via the Anthropic Message Batches API (scripts/batch-rescore.py, batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx). Blinding: structural. The scorer received one tool-less request containing the scoring brief, the anchored scales, the scoring disciplines, and the evidence packet below inlined as text. It had no filesystem access and no search, so it could not reach ws09-scorecard.md, ws09-red-team-log.md, or deviations-log.md by any route.
Brief (also archived): the batch workdir brief is reproduced under § Brief below; source path at collection time was /private/tmp/gubment-batch-rescore-gbmt11-12/disability/BRIEF.md.
Evidence packet manifest (inlined into the request; not re-committed as duplicates of disability/research/ / disability/docs/):
- 00-scales.md
- phase0-findings.md
- ws01-baseline.md
- ws02-adjudication.md
- ws03-integrity.md
- ws04-work-cliff.md
- ws05-ticket-vr.md
- ws06-childhood-ssi.md
- ws07-precedents.md
- ws08-political-economy-seam.md
What it did not receive. ws09-scorecard.md, ws09-red-team-log.md, deviations-log.md, any site/disability/*, and any prior cell values.
Reconciled cell by cell in ws09-rescore-log.md. What follows is the scorer's output, verbatim and unedited, then the brief.
GBMT-12 Disability / SSDI–SSI — blind re-score
Cell IDs: rows 1–10 in the order given (1 = a1_adjudication_surge … 10 = a10_do_nothing); columns A–E = O1, O2, O3, O4, O5.
1. The matrix
| Architecture | O1_measurement_integrity | O2_adjudication_timeliness_accuracy | O3_poverty_insurance_protection | O4_work_without_cliff | O5_administrability |
|---|---|---|---|---|---|
| a1_adjudication_surge | 3 | 4 | 3 | 3 | 4 |
| a2_alj_consistency | 4 | 4 | 3 | 3 | 4 |
| a3_cliff_redesign | 3 | 3 | 4 | 2 | 4 |
| a4_medicaid_buyin | 3 | 3 | 4 | 4 | 4 |
| a5_ticket_vr | 3 | 3 | 3 | 2 | 3 |
| a6_cdr_wage_integrity | 4 | 3 | 2 | 3 | 4 |
| a7_partial_disability | 3 | 3 | 3 | 3 | 2 |
| a8_childhood_ssi_track | 4 | 3 | 3 | 3 | 3 |
| a9_definition_tightening | 2 | 2 | 2 | 3 | 2 |
| a10_do_nothing | 3 | 3 | 3 | 2 | 3 |
2. Cell-by-cell basis
a1 — Adjudication capacity surge
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 1A | 3 | No evidence in the record that a capacity surge improves or harms measurement honesty of the integrity object. Phase 0 band (no clean point fraud rate) is inherited, not moved. I deliberately did not credit anchor-4 here: the O1 anchor's exemplars ("CDR cessations, wage/SGA reporting") describe integrity-framed series, and a1's claimed object is backlog, not integrity. | unevidenced-neutral |
| 1B | 4 | §2: hearing APT fell 450 (FY2023) → 342 (FY2024) with pending hearings down ~19% YoY (OIG MMC / APM); §2 verdict and Implications state "Architecture #1 (adjudication capacity surge) … stay[s] first-class for O2/O5." Not 5: 342 > SSA's own 270-day goal, and initial APT rose 218 → 231 with ~1.18M pending initials. Caveat I am flagging: the record does not causally attribute the FY2024 hearing improvement to staffing/hiring/digitization; I am resting partly on the project's own explicit O2 implication. | judgment-call (evidenced trend + project implication) |
| 1C | 3 | The record contains no poverty, deep-poverty, or coverage-continuity outcome series tied to wait length. §7's steelman prose mentions claimants "wait through hardship," but that was not one of its three scored legs and carries no measured protection outcome. Scored the wait evidence once, in O2, rather than twice. A reviewer who reads O3's anchor-2 parenthetical ("long unprotected waits") as a standalone trigger would move this to 4 on wait reduction; I flag that as the live divergence. | unevidenced-neutral |
| 1D | 3 | No cliff-related evidence for adjudication capacity. | unevidenced-neutral |
| 1E | 4 | §2: SSA drew hearing pending down ~60k in one year (322k → 262k) and cut hearing APT by 108 days inside existing agency tools — anchor-4 "executable rollout inside existing agency tools"; §2 Implications name #1 first-class for O5. Not 5: initial-level capacity worsened in the same year (APT 218→231, ~1.18M pending), so "success at material scale" is not clean. | evidenced |
a2 — ALJ consistency / quality controls
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 2A | 4 | §2: GAO-18-37 measures residual allowance dispersion after claimant/judge/office case-mix controls (up to 46pp, 5th–95th percentile) — a separable, correctly-specified administrative series; §2 explicitly rules out the mislabel this architecture would otherwise ride ("'ALJs are soft' talking points that cite the 51% hearing allow rate without case-mix fail this workstream"; exclusion of "raw ALJ league tables without case-mix"). The row spec ("without a blanket allowance crackdown") is the case-mix-honest version. Judgment call: anchor 4's exemplars are integrity/fraud-adjacent series, so applying it to residual-variance measurement is an extension of the anchor. | judgment-call |
| 2B | 4 | §2: residual judge variance fell ~5pp over 2007–2015, "SSA attributed the narrowing to quality assurance and training"; OIG A-12-17-50220 found focused reviews move some outliers. This is anchor-4 verbatim ("residual-variance narrowing under quality … interventions"). Not 5: ~5pp against a 46pp residual range is not a large reduction, and no wait-standard clearance. | evidenced |
| 2C | 3 | No protection/poverty/coverage outcome evidence for consistency controls. (§2's finding that representation nearly tripled allowance probability is a case-processing fact, not an effect of this instrument.) | unevidenced-neutral |
| 2D | 3 | No cliff evidence. | unevidenced-neutral |
| 2E | 4 | §2: the narrowing instrument (quality assurance, training, focused outlier reviews) was already run by SSA/OIG inside existing tools with a measured effect — anchor 4. | evidenced |
a3 — Cliff redesign (1 − for−2, simplified reporting, automatic 1619(b)/Medicare)
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 3A | 3 | KC3 does not fire (1619(b) take-up published) and BOND supplies clean measured series, but nothing in the record shows this architecture improving or damaging the integrity object's measurement. Band inherited. | unevidenced-neutral |
| 3B | 3 | No evidence on DDS/ALJ timeliness or accuracy from an offset or from automatic continuance. Reduced post-entitlement/overpayment workload is a mechanism story only. | unevidenced-neutral |
| 3C | 4 | §4/§5 (BOND FER, Gubits et al. 2018): the offset raised average SSDI benefits due (+143/yr ≈ +1450–500/yr ≈ +4% Stage 2) and raised the share with earnings above BYA (+7% / +23% relative) — measured retention of cash for people who would otherwise step to zero at SGA. Judgment call: the measured object is benefits paid, not a poverty rate; §4 characterises much of it as a windfall to those already at SGA. A reviewer scoring strictly on poverty outcomes would hold 3. | judgment-call |
| 3D | 2 | §4/§5: BOND found null average earnings impacts in both Stage 1 (nationally representative) and Stage 2; EWIC counselling enhancement produced "virtually no" incremental earnings; 1619(b) take-up is 2.6% of blind/disabled 18–64 while the DI cash cliff persists. Anchor 2 verbatim. Noted: the bundle's automatic-extension/simplification arms are unevidenced for effect (§4 calls simplification "live," not measured), so they do not lift the cell. | evidenced |
| 3E | 4 | §4/§5: SSA actually operated a 1 − for−2 offset with WIC/EWIC arms for 4–5 years across a national Stage 1 sample plus volunteer Stage 2 (Abt/Mathematica BOND FER) — anchor-4 partial evidence of executable operation inside agency tools; 1619(b)/extended Medicare are existing statutory machinery (§4 Red Book stack). Counterweight logged: BOND Stage 1 showed net social cost for the full caseload — a financing finding, not a missing-precondition or failed precedent, so it does not push to 2. | judgment-call |
a4 — Medicaid buy-in expansion
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 4A | 3 | §4/§7 document a real measurement gap: last comprehensive national buy-in enrollment is ~193,000 across 35 states in 2011 (BPC 2022 data-gap brief), explicitly called "a data gap for O4 scoring." But a stale series is not the mislabel/construct-mix the O1 anchors grade, so I did not push below 3. Anchor fit is poor here. | judgment-call (evidenced gap, anchor doesn't cover it) |
| 4B | 3 | No adjudication evidence. | unevidenced-neutral |
| 4C | 4 | §4: Washington MBI matched-comparison evaluation (Gettens et al. 2012, JDPS) — MBI entrants showed more work, higher earnings, and lower food-stamp reliance while preserving medical coverage vs matched conventional-Medicaid stayers. Anchor-4 coverage continuity. Not 5: §4 grades the design "moderate-to-strong," and the national MBI series is pre/post among enrollees. | evidenced |
| 4D | 4 | §4 H4 leg 2: WA MBI matched study + Mathematica MBI "three E's" (~40% of participants with wages increased earnings; median real increase ~$2,600), explicitly contrasted with the BOND EWIC null and Ticket ITT nulls — anchor 4 verbatim ("coverage-separation … moves earnings/work more than information-only tools"). §4 Implications: "#4 … is the cleaner O4 instrument on present evidence." Not 5: no population-scale causal estimate. | evidenced |
| 4E | 4 | §4: 47 states already offer a buy-in (KFF 2025/26 surveys) under existing BBA/Ticket-era state options — anchor-4 "executable rollout inside existing agency tools / state options." Not 5: enrollment last tallied at ~193k (2011) and §8 documents state Medicaid cost-shift/FMAP friction as a real pivot point. | evidenced |
a5 — Ticket / VR redesign
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 5A | 3 | §5 documents a construct-mix risk in Ticket discourse (participant-only STW rates ~19–29% offered as population impact), but the row spec requires "honest effect sizes," and no cited evidence shows this architecture installing or amplifying a mislabeled series. | unevidenced-neutral |
| 5B | 3 | No adjudication evidence. | unevidenced-neutral |
| 5C | 3 | No poverty/coverage evidence for Ticket in either direction. | unevidenced-neutral |
| 5D | 2 | §5: SSA-commissioned ITT evaluations (Stapleton, Mamun & Page; SSB v83n1; Wittenburg et al. 2007) find no consistent, significant increase in STW or months in STW; detected effects are service enrollment only, 0.1–0.4pp; undetectable residual bounded at ~5% relative ≈ ~0.025pp absolute; DI SGA-termination rates flat at 0.35–0.55% across the TTW era; assignment 1–5% of eligibles. Anchor 2 verbatim. | evidenced |
| 5E | 3 | Evidence both ways: the program has operated nationally for two decades inside existing tools (executable), but the outcome-payment channel it depends on barely materialised — EN service enrollment 0.04% of the 2006 stock cohort vs SVRA 1.2%, ~90%+ of in-use tickets still to SVRAs, EN+SVRA success statistically indistinguishable from SVRA-only, and Thornton's ~3,000-STW-cases/yr self-financing threshold not reached (§5). Nets to neutral. | evidenced-wash |
a6 — CDR / wage-reporting integrity
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 6A | 4 | §3: the integrity object here is exactly the separable series the O1 anchor names — CDR cessation person-counts (~39,056 disabled-worker cessations FY2019; 136,481 DDS cessations pre-COVID) and wage/SGA reporting (SGA overpayments 86% driven by untimely earnings reporting) — versus CDI judicial actions of 74–77/yr, two orders of magnitude smaller. §3 Implications: "#6 … outranks prosecution-theater frames on the evidence in this file." Not 5: no clean intentional-fraud series exists (KC4 does not fire). | evidenced |
| 6B | 3 | Real evidence both ways and no measured timeliness/accuracy outcome. CDR work is DDS work whose volume tracks funding/capacity (FY2019: 215,720 full medical reviews + 766,913 mailers; OIG A-01-21-51038 DDS workload audit shows cessations halving in the COVID window) inside a system with ~1.18M pending initials and rising initial APT — but no cited measurement links CDR volume to initial/hearing waits, and the wage-reporting arm plausibly reduces post-entitlement work. Flagged: a stricter reading of O2 anchor 2 ("documented mechanism that consumes scarce DDS/ALJ capacity … without a measured timeliness/accuracy gain") would score this 2. | evidenced-wash |
| 6C | 2 | §3/§7: the instrument's own output is benefit exit at scale — ~39k disabled-worker cessations + FO FTC terminations (FY2019) and ~136k DDS cessations in a pre-COVID year; §6 shows the same instrument class driving ~⅔ of a >25% child SSI caseload decline 2014–2021. That is documented cash/coverage loss attributable to the instrument, with no offsetting coverage gain in the record. Anchor-2 "documented partial adverse protection effect." Construct caution honoured: AFR text says cessation does not imply the original award was wrong, so this is a protection-outcome score, not a finding of wrongful termination. | evidenced |
| 6D | 3 | Mechanisms conflict and neither is measured: better wage reporting could reduce overpayment shocks around SGA; tighter SGA/wage enforcement could sharpen the cliff. No cited effect either way. | unevidenced-neutral |
| 6E | 4 | §7 §2.2: "When funded and staffed, they produce tens to hundreds of thousands of reviews and tens of thousands of cessations per year" — FY2019 volumes confirm operation at material scale inside existing agency tools. Held at 4 rather than 5 because the same record documents collapse when unfunded (child CDRs ~150k → ~45k FY2000–2011, GAO-12-497) and COVID-window halving. | evidenced |
a7 — Partial disability / graduated benefits
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 7A | 3 | No measurement-integrity evidence for graduated schedules in this record. | unevidenced-neutral |
| 7B | 3 | The record's adjudication-load caution for reassessment intensity (UK ESA/PIP) is assigned explicitly to architecture #9, not #7 (§7 §3.1); §7's demotion instruction for #7 names O5/O4 only. Honouring that scoping. | unevidenced-neutral |
| 7C | 3 | No protection-outcome evidence for partial schedules. VA pays partial awards without an SGA-incapacity test (CRS R41289; SSA EN-64-125) but no poverty/coverage effect is measured. | unevidenced-neutral |
| 7D | 3 | Genuine wash and a disclosed divergence point. Positive: the US already runs a large federal graduated schedule (VA, 0–100% in 10% increments) that is "work-compatible by design" and does not require inability to engage in SGA. Negative: §7 states importing a schedule "without rewriting the SGA cliff and Medicare/Medicaid link is not a drop-in reform," and the NL/Nordic models presuppose an employer-financed sickness stack and non-cliff coverage the US lacks. §7 literally says "demote them on O5/O4"; I discharged that demotion in O5 where the anchor language matches verbatim, and read the row spec's mandatory transferability caution as satisfying §7's "preconditions explicit" branch here. A reviewer applying the instruction literally to both axes would score 2. | evidenced-wash / judgment-call |
| 7E | 2 | §7 §3.2: Dutch WIA rests on employers financing up to two years of sickness benefit plus experience-rated contributions; the Nordic partial-benefit family sits on universal/employment-linked coverage and active labour-market institutions — "US SSDI lacks the employer-sickness mandate and the non-cliff health coverage stack those models presuppose." Anchor-2 "capacity or financing preconditions the record shows are missing or non-portable." Precedent clause partially met: §7's cautionary precedent (UK ESA/PIP turbulence, "not fit for purpose") is formally assigned to #9, and VA is a working domestic schedule precedent — which is why I did not go to 1. | evidenced |
a8 — Childhood SSI separate track
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 8A | 4 | §6: the child integrity object in the record is a separable administrative series — childhood CDR volumes/cessations (overall child CDRs ~150k → ~45k FY2000–2011; mental-impairment CDRs ~84k → ~16k; SSB v84n4 attributing ~⅔ of the 2014–2021 caseload decline to CDR cessation volume) and evidence completeness — and §6 explicitly forbids importing the adult stewardship IP percentage as a child fraud rate ("Importing §3's SSI improper-payment percentage … as a childhood fraud rate fails the construct test the same way adult IP≠fraud fails KC1"). Also holds the awards≠stock≠SSI lock (§8). Judgment call: this credits the architecture for refusing a documented mislabel, which is anchor-4-adjacent rather than anchor-4 exact. | judgment-call |
| 8B | 3 | GAO-12-497 documents the problem (school-evidence obstacles, inconsistent secondary-impairment data, ~54% denial rate for child mental claims, overdue CDR backlogs) but no evaluated instrument effect on child timeliness or accuracy. Size of problem ≠ evidence of effect. | unevidenced-neutral |
| 8C | 3 | Child SSI is need-tested cash (~983,169 recipients, ~$793/mo average federal payment) and §6 says caseloads move with poverty and CDR intensity — but no measured poverty/coverage effect of a separate-track design. | unevidenced-neutral |
| 8D | 3 | §6: "No work cliff as the binding adult object" — children are not in the TWP/EPE/SGA path. This dimension is closer to not-applicable than to neutral; the anchors fit badly and I did not manufacture either direction. | unevidenced-neutral (anchor misfit) |
| 8E | 3 | Evidence both ways: child CDR capacity has been executed at large scale post-2013 (driving ~⅔ of a >25% caseload decline) and has also collapsed when unfunded (150k→45k, with overdue-review backlogs), and DDS examiners report obstacles obtaining the school evidence the child standard requires (GAO-12-497). Nets to neutral. | evidenced-wash |
a9 — Definition tightening / listings reform
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 9A | 2 | The constructs this architecture's public case runs on are each documented as mislabeled and uncorrected in the record: the raw-roll "explosion" (§1 — age-sex-adjusted incidence 6.4/1,000 in 2010 → 2.9 in 2022–23 → 3.3 in 2024; "fraud is not a measured driver in either decomposition"; "O1 cells that rest on 'rolls exploded' must cite adjusted incidence"); the 51%-hearing-allowance-as-ALJ-softness claim (§2); and the $72B IP-relabeled-as-fraud citogenesis root (§8). §7 puts this architecture in the same demotion bucket as fraud-first frames via the KC4 re-rank path. Judgment call: the record never states in so many words that definition tightening depends on the mislabel; I inferred that from its "get tougher" framing plus §7's explicit joint demotion. Reviewers who require the literal statement should hold 3. | judgment-call |
| 9B | 2 | §2 verdict, verbatim: "Definition-tightening architectures that ignore DDS/ALJ capacity inherit O2/O5 penalties," in a system with ~1.18M pending initials, initial APT rising 218→231, hearing APT 342 vs a 270 goal. §7 §3.1: UK ESA/PIP reassessment intensity without adjudication capacity produced sustained tribunal overturns and repeated "not fit for purpose" process critiques. §7 leg (a): the front door is already at 38%/38.3% initial allowance. Anchor 2 ("consumes scarce DDS/ALJ capacity … without a measured timeliness/accuracy gain"). Not 1: the UK evidence is a foreign analogue, not direct US measurement. | evidenced |
| 9C | 2 | §7 §2.1: PRWORA is characterised as "a definition/eligibility cut, not a fraud-rate discovery," with measured caseload effects running through eligibility rules (child caseload ~955k in Dec 1996 → ~847k in Dec 2000; SSI eligibility on DA&A primary impairment ended); §7 §3.1 records claimant-harm narratives under UK reassessment waves. Anchor 2 names "eligibility cuts" explicitly. Not 1: §6/§7 also show rolls rebounding without statutory change, so a durable coverage-loss attribution is not clean. | evidenced |
| 9D | 3 | No cited evidence that definition tightening steepens or eases the earnings cliff. Tightening the front door is a different object from the SGA step-function; I declined to score the mechanism. | unevidenced-neutral |
| 9E | 2 | Same citations as 9B: §2's explicit O5 penalty for capacity-blind tightening; §7's UK ESA/PIP cautionary precedent (backlog, error, backlash) assigned expressly to "#9"; documented current DDS/ALJ strain. Anchor 2 verbatim, including the "cautionary failed/turbulent precedent" clause. | evidenced |
a10 — Do-nothing comparator
Treated identically to the other rows: no penalty for being a comparator, no charity.
| Cell | Score | Basis | Flag |
|---|---|---|---|
| 10A | 3 | Anchor 3's second clause applies directly: the cell inherits the band that there is no clean point intentional-fraud rate (Phase 0 KC1 fires / KC4 does not fire; §3 finds no administrative series measuring intentional fraud at circulating magnitudes). Evidence on both sides: SSA's own AFR text labels the construct correctly ("stewardship reviews of the nonmedical aspects"; "we consider the benefits received before improvement to be proper"), while §8 shows the same cumulative being relabeled as fraud downstream. I did not score below 3 because the mislabel in the record is downstream of SSA's reporting, not inside it. | evidenced-wash |
| 10B | 3 | Genuinely mixed in the record: hearing APT improved 450→342 with pending down 19%, while initial APT lengthened 218→231 with ~1.18M pending initials, hearing waits stayed above the 270-day goal for a multi-year span, GAO's 46pp residual judge dispersion persists, and AC remands rose to 18.05% in FY2024 (then 15.54% FY2025) with 63% federal-court remands. Direct evidence of both lengthening and shortening → wash, not a downgrade. | evidenced-wash |
| 10C | 3 | The record contains no poverty or deep-poverty series at all, and no coverage-continuity outcome for applicants during the wait. Documented long waits (231/342 days) sit against the status quo's actual delivery of benefits to 7,365,987 disabled workers and 6,264,542 blind/disabled SSI recipients. Disclosed divergence: reading O3 anchor 2's "long unprotected waits" as a standalone trigger would score this 2; I held 3 to avoid scoring the same wait evidence twice (O2 and O3) and because no protection outcome is measured. | evidenced-wash / judgment-call |
| 10D | 2 | Anchor 2's second clause verbatim — "take-up stays negligible while the cliff remains": §4 Red Book stack shows the cash step to zero at persistent SGA after TWP/grace; 1619(b) participation is 2.6% of blind/disabled 18–64 (108,825 of 4,148,143); working blind/disabled 5.1%; Ticket assignment ≤~5% of eligibles with ITT nulls; DI termination-for-SGA 0.35–0.55%/yr of the worker stock. | evidenced |
| 10E | 3 | Current operations do execute at material scale (2,086,885 initial decisions, 495,700 reconsiderations, 289,492 hearings in FY2024; hearing backlog drawn down 19%), while simultaneously missing SSA's own 270-day standard with ~1.18M pending initials and OIG carrying hearings as a standing major-management challenge. Evidence both ways → 3, not 4. | evidenced-wash |
3. What my matrix implies
On my own numbers, a2 (ALJ consistency/quality controls) and a4 (Medicaid buy-in expansion) come out strongest, for opposite reasons: a2 is the only architecture with a measured improvement in its own core dimension (GAO's ~5pp residual-variance narrowing under quality assurance/training) plus a clean, case-mix-honest measurement object; a4 is the only architecture with cited positive evidence on both protection and the work cliff (WA MBI matched study; Mathematica MBI earnings series) while already operating as a state option in 47 states. a1 trails them narrowly on documented wait/backlog movement and executability. a9 (definition tightening) is clearly weakest and is the only row I scored below neutral on four of five axes — it is the one architecture the record penalises explicitly and by axis (§2 "inherit O2/O5 penalties"; §7 UK cautionary precedent; §7 KC4 demotion). Otherwise the board is mostly flat: 28 of 50 cells sit at 3, nine of those because real evidence washes out and nineteen because the record simply has nothing on that cell. Three rows tie at 14 (a5, a7, a10) by very different routes, and a3 and a10 are identical on O4 (both 2) — I am saying that plainly: BOND's null plus 2.6% 1619(b) take-up means the record does not distinguish a cash-offset cliff redesign from the status quo on the work-without-a-cliff axis. The O3 column is the weakest-evidenced column on the board; the O1 column is largely a band by construction.
4. Evidence-coverage count
Cells scored on cited evidence (non-3), plus evidenced-wash 3s counted separately:
| Architecture | Non-3 cells (cited evidence of effect/failure) | Additional evidenced-wash 3s | Unevidenced-neutral 3s |
|---|---|---|---|
| a1_adjudication_surge | 2 (O2, O5) | 0 | 3 |
| a2_alj_consistency | 3 (O1, O2, O5) | 0 | 2 |
| a3_cliff_redesign | 3 (O3, O4, O5) | 0 | 2 |
| a4_medicaid_buyin | 3 (O3, O4, O5) | 0 (O1 = judgment-call 3) | 1 |
| a5_ticket_vr | 1 (O4) | 1 (O5) | 3 |
| a6_cdr_wage_integrity | 3 (O1, O3, O5) | 1 (O2) | 1 |
| a7_partial_disability | 1 (O5) | 1 (O4) | 3 |
| a8_childhood_ssi_track | 1 (O1) | 1 (O5) | 3 |
| a9_definition_tightening | 4 (O1, O2, O3, O5) | 0 | 1 |
| a10_do_nothing | 1 (O4) | 4 (O1, O2, O3, O5) | 0 |
Totals: 22 non-3 cells, 9 evidenced-wash 3s, 19 unevidenced-neutral 3s.
5. What I could not score, and what would change my mind
Structural gaps in the evidence base.
- O3 has almost no outcome data. The record contains no poverty rate, deep-poverty rate, uninsurance rate, or hardship-during-wait measure anywhere. Everything I scored on O3 is either a program-parameter fact (benefits due, coverage preserved, cessation counts) or an eligibility-cut precedent. Four of ten O3 cells are unevidenced-neutral for this reason. A single cited poverty/coverage-outcome study — for wait length, for buy-in enrollees, or for CDR-ceased beneficiaries — would move 1C, 8C, 10C, and possibly 6C and 7C.
- O1 is a band by design. Phase 0 instructs every O1 cell to inherit a band rather than a point. That makes O1 largely non-discriminating except where an architecture's object is a separable administrative series (a6, a8) or is explicitly a mislabeled construct (a9). If §3's owed "separable intentional-fraud $ or person counts" extract ever lands (KC4), a6 could go to 5 and a9's O1 basis would need re-examination.
- No causal attribution for the FY2024 wait improvement. 1B and 1E rest on a documented improvement whose cause the record does not identify, plus §2's own explicit implication statement. Evidence that the 450→342 drop was driven by hiring/digitization would firm 1B toward 5; evidence it was driven by receipt declines or dismissal-rate changes (hearings show 33% dismissals) would pull 1B and 1E to 3.
- a7's O4 is my most contested call. §7 says "demote them on O5/O4"; I demoted only O5, where anchor 2's language matches word-for-word, and treated the row spec's mandatory transferability caution as satisfying §7's "preconditions explicit" branch on O4. If the reconciler reads the instruction as unconditional, 7D becomes 2. I would change it on any cited work/earnings outcome from a graduated-schedule system, in either direction — the record has none.
- a6's O2 is my second most contested call. The record establishes that CDRs are DDS work whose volume is funding-dependent, and establishes DDS scarcity, but never links CDR volume to initial-claim waits. If the owed OIG DDS-workload extract shows CDR waves displacing initial-claim processing, 6B goes to 2.
- a9's O1 rests on inference. The record documents the mislabels (raw rolls, hearing-vs-initial allowance, IP-as-fraud) and demotes definition tightening alongside fraud-first frames, but never states that definition tightening depends on those constructs. That is the one below-neutral cell on this board I would concede fastest to a 3.
- Cells the anchors fit badly. 8D (children are outside the work-cliff path entirely — closer to N/A than neutral); 4A (a documented stale-data gap that the O1 anchors, written around mislabeling, do not grade); 10E and 1E (an "administrability" scale whose anchors assume a rollout, applied to a do-nothing row and to a capacity row whose two stages moved in opposite directions).
- Not scoreable at all from this base: any child-SSI instrument effect (§6 supplies diagnosis, not evaluation); any effect of file digitization or ALJ hiring specifically; any effect of "simplified earnings reporting" as distinct from BOND's offset; current national Medicaid buy-in enrollment (last figure 2011); and Ticket's O5 under a redesigned payment rule as opposed to the one that has run.
Brief (verbatim)
GBMT-12 Disability / SSDI–SSI — blind re-score brief
Score all 10 rows × 5 objectives = 50 cells. Use only the evidence base appended to this message. Do not invent facts outside it. "The evidence base does not say" is a valid finding and usually means score 3 (unevidenced-neutral), per the disciplines below.
Architectures (rows) — score as specified, not as an idealized redesign
Use these IDs exactly in your matrix:
| ID | Architecture | Spec to score |
|---|---|---|
| a1_adjudication_surge | Adjudication capacity surge | DDS staffing, ALJ hiring, file digitization; clear the backlog before tightening standards. |
| a2_alj_consistency | ALJ consistency / quality controls | Reduce unexplained variance without a blanket allowance crackdown. |
| a3_cliff_redesign | Cliff redesign | 1 − for−2 offsets, simplified earnings reporting, automatic 1619(b)/Medicare extensions. Score against BOND and take-up evidence — do not invent earnings gains BOND did not find. |
| a4_medicaid_buyin | Medicaid buy-in expansion tied to disability work attempts | Separate health coverage from cash exit. Score existing state buy-in class evidence. |
| a5_ticket_vr | Ticket / VR redesign or replacement | Outcome-based payments with honest effect sizes from Ticket/VR evaluations. |
| a6_cdr_wage_integrity | Integrity focused on CDRs and wage reporting, not fraud theater | Align integrity work with separable administrative series (CDR cessations, wage/SGA reporting). Improper-payment % is not intentional fraud in this record. |
| a7_partial_disability | Partial disability / graduated benefits | VA-like or Nordic-inspired schedules; transferability caution mandatory. |
| a8_childhood_ssi_track | Childhood SSI separate track | Distinct evidence base and instruments from adult DI. |
| a9_definition_tightening | Definition tightening / listings reform | The “get tougher” architecture — score on adjudication load, protection, and administrability jointly. |
| a10_do_nothing | Do-nothing comparator | Current backlog trajectory, current cliffs, current Payment Integrity reporting. Identical treatment; no charity, no penalty for being a comparator. |
Dimensions (columns) — use these exact names
Anchored 1–5 meanings are in 00-scales.md in the evidence base. Columns in matrix order:
O1_measurement_integrityO2_adjudication_timeliness_accuracyO3_poverty_insurance_protectionO4_work_without_cliffO5_administrability
Disability-specific notes (do not override the disciplines)
- SSDI ≠ SSI. Never treat them as interchangeable in a cell.
- Awards ≠ stock. Mental disorders are ~12.7% of recent disabled-worker awards and a larger share of stock — do not mix constructs.
- Improper payment ≠ intentional fraud. Stewardship IP rates are not a fraud series. Below-neutral O1 needs evidence the architecture depends on or amplifies that mislabel — not mere absence of a clean fraud rate.
- BOND found null mean earnings under a 1 − for−2 offset; do not score cliff cash-offset arms as if they cleared a large earnings bar.
- Ticket population employment/exit effects are null or undetectable (≪2pp) in the record’s ITT evaluations.
- Front-door initial allowance is already in the ~30–40% band in recent tables; “get tougher” competes with the same DDS/ALJ capacity that binds backlog.
- Prefer cited evaluation effects over mechanism stories. Mechanism without evidence → 3.