GUBMENTPlain talk · policy frontier
Filings / Disability / Sources / Structurally blinded re-score & reconc
GBMT-12 · Research record · No. 12

Structurally blinded re-score & reconciliation log (GBMT-12)

disability/research/ws09-rescore-log.md
This is a working research document from the disability filing, published as written — including the parts later corrected. It is the underlying record for Whitepaper No. 12, not a summary of it.

Date: 2026-08-11 Method: Per M6/M10 — structural blinding via scripts/batch-rescore.py (batch msgbatch_013vUDQ6mt6JV8mz2H95zfsx, model claude-opus-5). Raw scorer output, brief, and evidence-packet manifest: ws09-blind-scores-2026-08-11.md. Prior board: pass-1 + M6 red-team (ws09-red-team-log.md). This is the owed structurally blinded pass.

Scale posture (already corrected at pass-1). Disability’s published anchors already encode the elder-care / housing evidence-floor lesson: below-neutral requires cited adverse/undermining evidence; absence → 3; above-neutral requires cited evidence of the measured property. The blind brief used the same disciplines. Divergence here is almost entirely application of that floor, not a scale rewrite.


Headline

The blind scorer matched the published post-red-team board on 42 of 50 cells and differed on 8. Reconciliation moves 6 cells and logs 2 judgment splits (published value kept). Net effect:

Weighting Post-red-team top Pass-2 reconciled top
Applicant in the backlog #1 = #2 (3.60) #2 ALJ QC alone (3.70); #1 second (3.60)
Beneficiary who might work #4 (3.75) #4 unchanged (3.75)
Integrity hawk / budget #6 = #8 (3.60) #2 ALJ QC alone (3.80); #8 second (3.60); #6 drops to the 3.50 tie band

#9 “last under all three”: red-team withdrew that claim because #10 was last under beneficiary after O1 1→2. Pass-2 reinstates the arithmetic — not by cutting #9 further (its row is unchanged), but by lifting #10’s unevidenced below-neutral cells to 3. Under the reconciled board #9 is again last under every weighting (2.05 / 2.40 / 2.10). State that as a pass-2 consequence of the evidence floor on do-nothing, not as a quiet resurrection of the withdrawn red-team overclaim.

Rank-stable top set (top four under every weighting): #1 / #2 / #4 / #8. Post-red-team’s #1/#4/#6 set breaks: #6’s O3 3→2 exits the applicant top four; #2’s O1 3→4 enters every reader’s top band.


Reconciliation

Every differing cell is accounted for. “Resolved” is the pass-2 board value.

A. Adopt blind — elder-care floor (below-neutral unevidenced → 3)

Cell Published Blind Resolved Basis
#7 O2 2 3 3 Published bundled UK ESA/PIP reassessment turbulence into #7’s O2. The record assigns that caution expressly to #9 (ws07 §3.1); #7’s demotion instruction is O5/O4 only (§3.2: “demote them on O5/O4”). No cited O2 adverse for graduated schedules → floor returns the cell to 3. Classic elder-care correction.
#10 O1 2 3 3 Published treated live IP→fraud citogenesis as status-quo O1=2. Blind: SSA’s own AFR labels stewardship correctly; the mislabel is downstream of agency reporting (ws03, ws08). Anchor 2 requires the architecture rely on a contested construct — do-nothing does not operate on the mislabel. Absence of a clean fraud series is the KC1 band (anchor 3), not a free 2.
#10 O2 2 3 3 Published read “still above 270 / initial APT rising” as O2=2. Blind: hearing APT improved 450→342 with pending −19% while initials worsened 218→231 (ws02) — genuine mixed trajectory → evidenced wash at 3, not a one-sided adverse. Anchor 1 (lengthening) is not met board-wide; anchor 2’s “consumes capacity without gain” does not describe the status quo’s own hearing drawdown.

B. Adopt blind — cited evidence moves the cell

Cell Published Blind Resolved Basis
#2 O1 3 4 4 GAO-18-37 residual allowance dispersion after case-mix controls, plus SSA-attributed ~5 pp narrowing under QA/training (ws02), is a separable, correctly-specified administrative series. Row #2’s public case is unexplained variance without raw-allowance mislabel — construct alignment parallel to #6/#8’s O1=4 wins. Anchor-4 exemplars are integrity-framed; applying them to residual-variance measurement is an extension, but the measured object matches the architecture. Held at 4 not 5: residual still large (~46 pp).
#3 O5 3 4 4 SSA actually operated a national 1 − for2 offset with WIC/EWIC arms for years (BOND FER; ws04); 1619(b)/extended Medicare are existing statutory machinery. Anchor 4 (“executable rollout inside existing agency tools”) is met. Net social cost is a financing finding, not a missing-precondition → does not push to 2. Not 5: Stage 1 full-caseload cost and volunteer Stage 2 limit “success at material scale.”
#6 O3 3 2 2 The instrument’s measured output is benefit exit at scale — ~39k disabled-worker cessations (FY2019) / ~136k DDS cessations in a pre-COVID year (ws03); child CDR intensity drives ~⅔ of a >25% caseload decline (ws06). Anchor 2 (“documented partial adverse protection effect”) fits. Construct caution preserved: cessation ≠ original-award fraud — this is a protection-outcome score, not a wrongful-termination finding.

C. Judgment splits — published value kept; blind reading recorded

Cell Published Blind Kept Why
#3 O3 3 4 3 Blind credits BOND’s rise in average benefits due (+143/+450–500) as protection. ws04 characterises much of that as a windfall to those already at SGA, not a poverty/coverage outcome. Red-team Attacks 2/4 already barred above-neutral O3 on mechanism or non-poverty proxies. Anchor 4 wants poverty reduction or coverage continuity — benefits-due windfall is neither. Blind itself flagged the poverty-strict reading as the live divergence.
#8 O5 4 3 4 Blind nets child CDR capacity boom/bust to an evidenced wash. The architecture under score is a separate child track, and PRWORA already executed a child-specific standard at national scale (ws06, ws07) — that is anchor-4 administrability for the row as specified. CDR funding volatility is real but is #6’s operating fact, not a veto of child-track executability.

D. Unchanged — blind independently matched published (42 cells)

#1 all five; #2 O2/O3/O4/O5; #3 O1/O2/O4; #4 all five; #5 all five; #6 O1/O2/O4/O5; #7 O1/O3/O4/O5; #8 O1/O2/O3/O4; #9 all five; #10 O3/O4/O5.

Material matches worth naming: #3 O4=2 (BOND null) and #5 O4=2 (Ticket ITT ≪2pp) re-derived independently; #9’s four below-neutral cells and O4=3 matched exactly (including post-red-team O1=2); #4’s 3/3/4/4/4 buy-in row matched cell-for-cell; #1’s post-red-team O3=3 held without re-inflation.


Reconciled matrix

# Architecture O1 O2 O3 O4 O5
1 Adjudication capacity surge 3 4 3 3 4
2 ALJ consistency / QC 4 4 3 3 4
3 Cliff redesign 3 3 3 2 4
4 Medicaid buy-in expansion 3 3 4 4 4
5 Ticket / VR redesign 3 3 3 2 3
6 CDR / wage integrity 4 3 2 3 4
7 Partial disability 3 3 3 3 2
8 Childhood SSI separate track 4 3 3 3 4
9 Definition tightening 2 2 2 3 2
10 Do-nothing 3 3 3 2 3

Bold = moved vs post-red-team published board.

Recomputed weightings

Rank Applicant in backlog Score Beneficiary who might work Score Integrity hawk / budget Score
1 #2 ALJ QC 3.70 #4 Medicaid buy-in 3.75 #2 ALJ QC 3.80
2 #1 Capacity surge 3.60 #2 ALJ QC 3.35 #8 Childhood track 3.60
3 #4 Medicaid buy-in 3.50 #8 Childhood track 3.25 #1 = #4 = #6 3.50
4 #8 Childhood track 3.30 #1 Capacity surge 3.20
5 #3 Cliff redesign 3.15 #6 CDR/wage 3.00
6 #5 Ticket = #10 Do-nothing 2.95 #7 Partial disability 2.90 #3 Cliff 3.20
7 #3 Cliff 2.70 #5 Ticket = #10 Do-nothing 2.90
8 #7 Partial disability 2.80 #5 Ticket = #10 Do-nothing 2.60 #7 Partial 2.70
9
10 #9 Definition tightening 2.05 #9 Definition tightening 2.40 #9 Definition tightening 2.10

Unweighted means (not a reader): #2 = #4 = 3.60 · #1 = #8 = 3.40 · #6 = 3.20 · #3 = 3.00 · #5 = #7 = #10 = 2.80 · #9 = 2.20.

Claims withdrawn or narrowed

  1. Integrity-hawk co-lead #6 = #8 — withdrawn as a tie headline. #2 leads alone after O1 lift; #6’s O3 demotion drops it into the 3.50 band with #1/#4. #8 remains second on integrity hawk as a construct artifact (honesty box retained).
  2. Applicant #1 = #2 co-lead — narrowed; #2 leads alone (3.70 vs 3.60).
  3. “#10 last under beneficiary” (the red-team replacement for “#9 last under all three”) — withdrawn. Do-nothing rises to 2.60 under beneficiary (ties #5); #9 is last under all three by pass-2 arithmetic.
  4. #9 last under all three — not reasserted as the red-team’s withdrawn overclaim. Stated as: under the reconciled floor-corrected board, #9 is last under every weighting because #10’s unevidenced 2s were returned to 3, while #9’s four evidenced below-neutral cells were independently re-derived.

What survives unchanged


Artifacts

← All Disability research documents Sources digest Read the whitepaper