GUBMENTPlain talk · policy frontier
Filings / Disability / Sources / Research Inquiry: Disability / SSDI–SS
GBMT-12 · Research record · No. 12

Research Inquiry: Disability / SSDI–SSI (GBMT-12)

disability/docs/research-inquiry.md
This is a working research document from the disability filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 12, not a summary of it.

Status: Protocol drafted 2026-08-09. Phase 0 landed 2026-08-11. Workstreams §1–§8 complete. Scorecard pass 2 (structurally blinded re-score reconciled) landed 2026-08-11 (research/ws09-scorecard.md, research/ws09-rescore-log.md, research/ws09-blind-scores-2026-08-11.md); prior red-team in research/ws09-red-team-log.md. Starred anchors 1–8 verified in Phase 0; KC1 fires; KC2 and KC3 do not. H1–H7 SUPPORTED; H8 SUPPORTED-WITH-DEVIATION (awards ≥25% fails at 12.7%; stock/SSI legs carry). Anchors 9–11 verified; 10 and 12 directionally verified (CDR≫prosecutions; rep-payee mental concentration). Pass-2 tops: backlog → #2 ALJ QC; might-work → #4 Medicaid buy-in; integrity-hawk → #2 ALJ QC (#8 second as construct artifact). Rank-stable top set #1/#2/#4/#8. #9 last under all three on pass-2 arithmetic (do-nothing floor lift); fraud-theater ghost demoted below #6. Designation: Numbered filing GBMT-12. Directory disability/. Branch gbmt-12-disability per M9 after Phase 0 lands; claim with scripts/preflight-filing.sh disability and git fetch before any substantive work. Scope: United States; the federal disability cash system — SSDI and SSI as administered: who is on the rolls and why the counts move, how applications are adjudicated (DDS → reconsideration → ALJ), what "fraud" and improper-payment figures actually measure, and whether a beneficiary can work without falling off a cliff (cash + Medicaid/Medicare interaction). Method: Imports method/gubment-method.md (M1–M10) in full. This document contains only domain content. Proportionality (M10): Three headline findings maximum, at full M1 rigor. Supporting claims: one typed, dated source, root-traced where citogenesis is plausible. Scorecard cells: cited to a workstream finding or held neutral, no separate verification pass. Target ≤8 findings files. One red team, one structurally blinded re-score, reconciliation as a table, deviations log at one line per entry. Artifact: A feasibility assessment of the leading "reform disability" architectures, naming binding constraints in priority order — Whitepaper No. 12.


1. Scope and definitional decisions to settle before fieldwork

1.1 The public fight, and the desk wager. The public disability argument is a moral binary — fraud vs. cruelty — hung on roll counts and an improper-payment percentage. This filing's wager is that the binding constraints are adjudication capacity (DDS/ALJ backlog and accuracy) and the work cliff (cash benefits interacting with Medicaid/Medicare and earnings rules), not the statutory definition of disability alone and not a fraud rate that may be single-root or definitionally non-comparable to "fraud." That wager is falsifiable; H2/H3 and KC1 are written to falsify it.

1.2 Which disability problem. At least five travel under the phrase: (a) roll size and composition — SSDI vs. SSI, age, diagnosis, concurrent benefits; (b) adjudication — initial allowance rates, appeals, ALJ variance, wait times; (c) integrity — improper payments, CDR continuance, fraud prosecutions vs. agency error; (d) work and the cliff — trial work period, Ticket to Work, substantial gainful activity, Medicaid 1619(b) / Medicare extensions; (e) childhood SSI — a distinct eligibility and political economy. The public argument collapses (a) and (c). This filing treats (b) and (d) as load-bearing and scores (c) only after the measurement root is audited.

1.3 SSDI ≠ SSI. SSDI is social-insurance (work credits, higher average benefits, Medicare after the waiting period). SSI is means-tested (income/asset rules, often Medicaid from day one, state supplements). Mixing them in a single "disability rolls" chart without labels is a finding about the discourse, not a baseline. Every workstream states which program it is analyzing.

1.4 Shared seam with GBMT-11 (Mental health). Mental disorders are a large share of working-age awards. This filing owns cash, adjudication, and cliffs; GBMT-11 owns care delivery and capacity. The shared seam — psychiatric award shares, representative-payee patterns, care scarcity as overflow into awards — is treated in §8 of each protocol. Neither filing re-litigates the other's object.

1.5 Excluded, with reasons.

1.6 Administrative geography. Initial disability determinations are made by state Disability Determination Services under federal rules and SSA funding — a federalist hinge. ALJ hearings are federal. Appeals Council and federal court review sit above. Which layer binds backlog and allowance variance is empirical (§2).

1.7 Perishability. Allowance rates, backlog counts, and improper-payment reports move with SSA budget, hiring, and CDR policy. The 2025–26 federal workforce and appropriations environment may disrupt series continuity. Every finding carries an as-of date; series breaks are reported, not smoothed.


2. Objective function

Candidate architectures are scored against five objectives:

# Objective What it measures
O1 Measurement integrity Whether the roll counts, award rates, and integrity statistics in the public argument measure what speakers claim
O2 Adjudication timeliness and accuracy Time to initial decision and hearing; remand/reversal patterns; ALJ variance beyond case mix
O3 Poverty and insurance protection Poverty rates, deep poverty, and health-coverage continuity for beneficiaries and denied applicants
O4 Work without a cliff Ability to attempt work without losing cash and Medicaid/Medicare in a single step; take-up of existing work incentives
O5 Administrability Whether SSA and DDSs can execute the instrument at current or plausible capacity — new rules, new determinations, lead time

O2 and O5 are in tension with "get tougher screening" architectures; O4 is in tension with "strict SGA cliff" integrity theories. The scorecard is expected to show the trades, not average them away.


3. Three objective weightings (M10)

Weighting The reader it represents Emphasis
The applicant in the backlog Waiting on DDS or an ALJ date; life on hold O2 heavy; O3 second; O5 as the means
The beneficiary who might work Wants earnings without losing Medicaid or facing a one-way exit O4 heavy; O3 second; O1 as honesty about rules
The integrity hawk / budget staffer Convinced the rolls are bloated or the agency cannot police them O1 and O5 heavy; O2 as quality control

4. Workstreams (8)

Seeds are seeds. Each workstream opens with a documented PRISMA-lite search (OpenAlex/Semantic Scholar; SSA.gov statistical publications; OIG/GAO; NASI) with strings, date ranges, and inclusion criteria logged. Findings files target: ≤8 (M10).

§1 Baseline — rolls, awards, and composition. Question: who is on SSDI and SSI, how did the counts move, and what share of the "explosion" is population aging, women's labor-force history, and diagnostic composition vs. policy looseness? Evidence: SSA Annual Statistical Report on the Social Security Disability Insurance Program; SSI Annual Statistical Report; ORES worker tables; age-adjusted award rates. Two-method rule on growth decomposition: SSA actuarial / published decompositions against an independent demographic reconstruction. Seed: H1.

§2 Adjudication backlog and accuracy (flagship). Question: where do cases sit (DDS vs. hearing), how long, at what allowance rate by stage, and how much ALJ variance remains after case mix? Evidence: SSA workload reports; hearing office statistics; GAO backlog reports; published ALJ allowance variance studies; Appeals Council and federal-court remand rates. Seed: H2. Load-bearing for O2/O5.

§3 The fraud / improper-payment number (kill-condition workstream). Question: what does the circulating integrity figure measure — fraud, agency error, change in beneficiary status, sampling artifact — and how many independent roots does it have? Evidence: SSA OIG improper-payment and cooperative-disability-investigations reports; Payment Integrity / AFR notes; CDR continuance statistics; criminal fraud prosecution counts vs. improper-payment dollars. Seed: H3. KC1 test.

§4 The work cliff and Medicaid/Medicare link. Question: which rules bind when a beneficiary earns — SGA, TWP, EPE, SSI income deeming, 1619(b), Medicare extended coverage — and what do take-up numbers say about usability? Evidence: SSA Red Book; 1619(a)/(b) participation; Ticket to Work / VR payment statistics; NASEM and academic work-incentive evaluations; state Medicaid buy-in programs. Seed: H4.

§5 Ticket to Work, VR, and measured return-to-work. Question: which return-to-work instruments have effects that survive independent evaluation, and how large relative to the cliff? Evidence: Ticket to Work evaluations (Mathematica / SSA); state VR outcomes; Benefit Offset National Demonstration (BOND) and successors; foreign partial-benefit experiments only with transferability notes. Seed: H5.

§6 Childhood SSI. Question: how large is the child SSI caseload, what drove post-1996 and post-pandemic moves, and which integrity/access claims are adult-program imports? Evidence: SSI child statistical tables; welfare-reform (PRWORA) child disability changes; school and Medicaid interactions; GAO child SSI reports. Seed: H6.

§7 Precedents. Question: what happened when this was tried? Domestic: PRWORA's disability-adjacent changes; CDR intensification episodes and their measured effects on exits vs. error; VA schedule-for-rating as a partial-disability precedent; state Medicaid buy-ins. International: UK ESA/PIP reassessment turbulence (cautionary); Netherlands / Nordic partial-disability and rehabilitation-first models with financing preconditions stated. Seed: H7 (steelman lives here).

§8 Political economy and the GBMT-11 seam. Question: who blocks cliff reform and who benefits from the fraud frame; what share of awards are mental disorders? Evidence: lobbying and member incentives on SSDI solvency vs. SSI appropriations; representative-payee data; SSA diagnostic tabulations (cross-cite GBMT-11); media citogenesis audit on one high-circulation fraud claim. Seed: H8.


5. Hypotheses with pre-registered adjudication criteria (M3)

H1 — The "explosion" is mostly composition and demography, not a mystery surge in fraud. Age-adjusted award rates and sex/composition shifts explain most of the long-run SSDI growth narrative still circulating. Supported if a published SSA or peer-reviewed decomposition attributes a majority of 1980s–2010s SSDI growth to aging, women's insured status, and related composition and age-adjusted award rates are flat or down in the most recent decade. Refuted if age-adjusted rates rise materially through the recent decade without composition explanation. Indeterminate if decompositions conflict across equal-quality sources — report the band.

H2 (flagship) — Adjudication capacity binds. Backlog and decision quality (including ALJ variance) constrain both access and integrity more than the statutory disability definition. Supported if hearing waits exceed a stated administrative standard in SSA's own reports for a multi-year span and ALJ allowance dispersion remains large after published case-mix controls in ≥1 major study. Refuted if waits are within standard and variance is largely explained by case mix. Indeterminate if SSA withholds hearing-level microdata needed for the variance test — report the opacity.

H3 (KC1 test) — The fraud number is not a fraud number. The headline improper-payment or "fraud" figure is dominated by definitional categories other than intentional fraud, or traces to a single root repeated as an administrative fact. Supported if the primary SSA/OIG improper-payment publication attributes a majority of dollars to error/complexity/status change rather than intentional fraud or every high-circulation "X% fraudulent" claim root-traces to one non-independent source. Refuted if an administrative series measures intentional fraud at the circulating rate with transparent methods. Indeterminate if publications use "improper" and "fraud" interchangeably without a separable breakout — that indistinguishability supports the weaker form of H3.

H4 — The cliff, not work aversion, binds return to work. Existing work incentives have low take-up because losing Medicaid/cash is the rational fear; simplifying the cliff moves earnings more than motivational programs alone. Supported if 1619(b)/Ticket take-up is low relative to the working-age rolls and ≥2 evaluations find cliff or offset designs move earnings more than information-only interventions. Refuted if take-up is high and earnings non-response persists under offset designs (BOND lineage). Indeterminate if take-up denominators are not published.

H5 — Ticket to Work is not the binding lever. Ticket/VR payment systems have small population-level effects relative to cliff redesign. Supported if official evaluations show employment or benefit-exit effects that are small relative to the beneficiary population (threshold: <2pp on employment for the eligible pool) in the primary SSA-commissioned evaluations. Refuted if effects exceed that bar in ≥2 independent evaluations. Indeterminate if only vendor or advocacy summaries exist without technical appendices.

H6 — Childhood SSI is a different program wearing the same logo. Adult integrity and work-cliff frames mis-travel when applied to child SSI without a separate evidence base. Supported if child caseload drivers (poverty, Medicaid interaction, school identification) differ in kind from adult SSDI drivers in SSA/GAO analyses and the major integrity interventions were designed on adult assumptions. Refuted if the same integrity metrics and work frames apply with minor modification. Indeterminate if child statistics lack the needed breakouts.

H7 — STEELMAN, unfashionable direction: the system is harsh at the front door and the integrity panic is backwards. Built with equal effort per M3. Initial allowance rates are low, many eventually-allowed claimants wait through hardship, CDRs already remove large numbers, and the fiscal risk that matters for SSDI is aging and earnings history — not street-level fraud. Supported if (a) initial allowance rates are below 40% in recent SSA tables, (b) a majority of hearing allowances are eventual allowances of the same claim rather than "new" disability, and (c) CDR cessations dwarf fraud prosecutions in person-count terms. Refuted if two of those three fail. Indeterminate if hearing-stage data cannot separate eventual-allowance from new claims.

H8 — Psychiatric awards are the seam, not the scandal. Mental disorders are a large share of working-age awards; treating that as proof of fraud without care-system evidence fails an independence audit (cross-cite GBMT-11). Supported if mental disorders are ≥25% of working-age SSDI or SSI awards in SSA tabulations and no independent administrative series shows fraud concentrated in that diagnostic group. Refuted if mental disorders are a small share or fraud is demonstrably concentrated there in primary data. Indeterminate if diagnostic tables are too coarse.

Tilt audit: H1–H6 and H8 lean toward "measurement and cliffs bind; fraud frame is overfitted." H7 is the counterweight and gets equal effort in §7.


6. Kill conditions (M6)

KC1 (fires as headline if confirmed) — integrity debate runs on an unmeasured or mislabeled number. If §3 confirms H3, the fraud-vs-cruelty argument is being conducted over a figure that does not measure intentional fraud (or is single-root). That is the filing's measurement headline, and every architecture scored on O1 inherits a band, not a point.

KC2 — backlog data are not stage-resolvable. If DDS vs. hearing waits and allowance rates cannot be reconstructed from public SSA publications for the verification year, O2 collapses to qualitative description and the filing reports the transparency failure.

KC3 — work-incentive take-up is unpublished. If 1619(b), Ticket, and related participation denominators are missing, O4 cannot be scored quantitatively; cliff architectures are ranked on rule logic and evaluation literature only, with the gap flagged.

KC4 (re-rank, not kill). If H3 is refuted — a clean intentional-fraud series exists at the circulating magnitude — §3 becomes a co-flagship with §2 and integrity architectures rise. If H7 is strongly supported, the whitepaper leads with front-door harshness and demotes fraud-first reforms before full scoring waste.


7. Anchor Table (M4)

Starred anchors 1–8 were verified in Phase 0 (2026-08-11). Non-starred rows remain unverified priors until workstreams. ★ = was Phase 0 priority.

# Anchor (unverified prior) Used in Verify against Verified value / delta
1 ★ SSDI disabled-worker beneficiaries ≈ 7.5–8.5M; SSI disabled recipients (incl. concurrent) on the order of 6–7M (recent year — do not sum casually) §1 SSA ASIDI / SSI Annual Statistical Report Mostly holds (Dec 2023): disabled workers 7,365,987 (just below prior floor); all DI disabled 8,709,006; SSI blind+disabled 6,264,542; concurrent 18–64 949,971. Do not sum casually.
2 ★ Initial DDS allowance rates ≈ 30–40% in recent years; hearing-level allowance rates materially higher §2, H2, H7 SSA outcomes reports; workload data Holds (FY2024): initial allow 38% (SAOR 38.3%); hearing allow 51%.
3 ★ Average wait for an ALJ hearing has been measured in hundreds of days in recent SSA/GAO reporting §2, H2 SSA hearing workload reports; GAO Holds. FY2024 hearing APT 342 days (FY2023 450); above 270-day goal. Initial APT 231 days. KC2 does not fire.
4 ★ Improper-payment rate for DI/SSI in Payment Integrity reporting is mid-to-high single digits or low double digits depending on program/year — and is mostly not labeled intentional fraud §3, H3, KC1 SSA AFR / PaymentAccuracy.gov; OIG Holds; KC1 fires. FY2023 SSI IP 10.62%; OASDI ~0.30%. Causes = reporting/complexity/status — not intentional fraud.
5 ★ Mental disorders ≈ 25–35% of working-age SSDI awards (diagnostic group tabulation) §8, H8, GBMT-11 seam SSA ASIDI diagnostic tables Corrected. Disabled-worker awards mental disorders 12.7% (2023), not 25–35%. Worker stock mental categories ≈ 28.6% (Dec 2023). Confirm GBMT-11 — do not reintroduce one-third-of-awards.
6 ★ Ticket to Work employment effects in SSA-commissioned evaluations are small at population scale §5, H5 Mathematica / SSA Ticket evaluations Holds; <2pp bar locked (§5). ITT STW/employment null or undetectable (≪2pp absolute); service enrollment +0.1–0.4pp (Wittenburg et al. 2007). H5 SUPPORTED.
7 ★ 1619(b) Medicaid continuance participation is a small fraction of working-age SSI recipients §4, H4 SSA SSI Annual Statistical Report; Red Book references Holds. Dec 2023: 108,825 1619(b) aged 18–64 = 2.6% of blind/disabled 18–64. KC3 does not fire. H4 take-up leg met.
8 ★ Age-adjusted SSDI incidence declined or flattened in the decade after the mid-2010s peak narrative §1, H1 SSA ORES / actuarial incidence tables Holds. Age-sex-adjusted incidence 6.4 (2010) → 4.0 (2019) → 2.9 (2022–23) → 3.3 (2024) per Trustees Fig. V.C3.
9 BOND (Benefit Offset National Demonstration) found limited earnings response to a 1 − for2 offset §4, §5 SSA/BOND final evaluation Verified (§4/§5). Stage 1/2: null mean earnings; Stage 1 benefits +143 * */yr(  + 1450–500/yr (~+4%); EWIC null vs WIC.
10 CDR cessations substantially exceed criminal fraud prosecution counts in person terms H7, §3 SSA CDR statistics; DOJ/SSA OIG prosecution summaries Directionally verified (§3/§7): FY2019 disabled-worker initial cessations+FTC ≈ 39k vs CDI judicial actions 74–77/year; DDS cessations can exceed 100k. H7(c) Met. Exact same-year paired CSV still polish-owed.
11 Child SSI recipients ≈ 1M order of magnitude in recent years §6 SSI Annual Statistical Report Verified (§6). Dec 2023 under-18 SSI 983,169; H6 SUPPORTED (PRWORA/GAO/CDR drivers ≠ adult DI).
12 Representative payees cover a large minority of beneficiaries with mental disorders §8, GBMT-11 seam SSA representative-payee reports Directionally verified (§8). Of ~1.7M disability beneficiaries with payees, 77.5% have mental disorders (review citing SSA); SSI ASR Table 37 shows heterogeneous payee rates by mental diagnosis (e.g. intellectual 67.9% vs depressive/bipolar 21.2% ages 18–64).

8. Candidate architectures to score

Scored per M6 against O1–O5 under the three weightings; each cell cited to a workstream finding or held neutral (M10).

  1. Adjudication capacity surge — DDS staffing, ALJ hiring, file digitization; clear the backlog before tightening standards.
  2. ALJ consistency / quality controls — reduce unexplained variance without a blanket allowance crackdown.
  3. Cliff redesign1 − for2 offsets, simplified earnings reporting, automatic 1619(b)/Medicare extensions; score against BOND.
  4. Medicaid buy-in expansion tied to disability work attempts — separate health coverage from cash exit.
  5. Ticket / VR redesign or replacement — outcome-based payments with honest effect sizes from §5.
  6. Integrity focused on CDRs and wage reporting, not fraud theater — if H3 holds, score this above prosecution-centric frames.
  7. Partial disability / graduated benefits — VA-like or Nordic-inspired schedules; transferability caution mandatory.
  8. Childhood SSI separate track — distinct evidence base and instruments from adult DI.
  9. Definition tightening / listings reform — the "get tougher" architecture, scored on O2/O3/O5 jointly (error cuts both ways).
  10. Do-nothing comparator — current backlog trajectory, current cliffs, current Payment Integrity reporting.

9. Phase 0 execution notes

Days, not months. The gate pass does five things:

  1. Verify starred anchors 1–8, with anchor 4 / KC1 first — read the improper-payment publication and separate fraud vs. improper vs. error in the agency's own words.
  2. Build the §1 baseline skeleton — SSDI and SSI counts with concurrent-benefit caution; age-adjusted incidence if published.
  3. Stage-resolve adjudication — initial vs. hearing waits and allowance rates for the verification year (KC2 test).
  4. Pull diagnostic award shares for the GBMT-11 seam (anchor 5) and representative-payee notes.
  5. Locate Ticket and BOND evaluation PDFs so §5 does not depend on press summaries.

Re-rank triggers. If KC1 fires, the whitepaper leads with measurement honesty and demotes fraud-first architectures. If H2 is weak (no backlog), cliff workstreams (§4–§5) become the flagship. If H7 is strongly supported, definition-tightening architectures are demoted early.

Freeze. This protocol at execution start is the commitment. Report structure, scorecard scales, and the three weightings are fixed before synthesis; deviations logged one line each — what changed, why, effect.

← All Disability research documents Sources digest Read the whitepaper