← Research Portal
AAL-D-006Published

Collateral Eligibility & Substitution

Benchmarks·Aug 2026·v1.1
250 (162 ineligible / 88 traps)
Cases
12 (ELIG-*)
Eligibility categories
9
Models evaluated
93.1% – 100%
Detection (range)
57.8% (Gemini 3.1 Pro) / 15.4% (Grok 4.5)
Value extraction (best / worst)
5.6% – 98.8%
Escalation (spread)
98.2% (Grok 4.5)
Substitution validity (best)

AAL-D-006 tests whether frontier models can adjudicate collateral eligibility and substitution: read a collateral schedule against an eligibility rulebook, decide what's ineligible, attribute the reason, value the exposure, propose a valid substitute, and decide whether to escalate. 250 cases (162 ineligible across 12 categories, 88 look-ineligible traps) across nine models — GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.5, DeepSeek V4 Pro, Kimi K3, Meta's Muse Spark 1.2, and Alibaba's Qwen 3.8-Max — three runs each, deterministic scoring, no model permitted to compute. Detection is universally solved (93.1–100%); value extraction tops out at 57.8% (Gemini 3.1 Pro) and bottoms at 15.4% (Grok 4.5). Two sharper findings: escalation accuracy spans 5.6% to 98.8% — wider than the value gap — and no single model wins across the board (Gemini 3.1 Pro leads on value, Opus 5 on escalation, Grok 4.5 on substitution — yet each hits the value wall). A separate cut (AAL-F-008) tests Meta's open-weight Muse Glimmer 30B against the same set — it matches the frontier field on detection and still hits the same value wall, evidence the quantification gap is architectural, not a function of scale.

Detection is solved; quantification is the wall

Every model clears 93% detection; none exceeds 58% on the actual dollar figure. A sixth consecutive workflow where knowing an exception exists is a solved language problem and computing its value is not.

Escalation, not value, splits the field

Escalation accuracy runs from 5.6% (DeepSeek V4 Pro) to 98.8% (Claude Opus 5) — a wider spread than value, and the metric that most directly governs production trust. A model that flags the ineligible piece but can't reliably decide when to escalate is a liability, not a control.

Ninth flagship, same wall

Meta's Muse Spark 1.2 and Alibaba's Qwen 3.8-Max join the roster at 100% detection each. Spark 1.2 posts the field's second-best substitution validity (92.6%) and still pins value accuracy at 31.1%. Qwen 3.8-Max does better — 39.2% value accuracy, second in the entire corpus behind only Gemini 3.1 Pro — but is also the first model in the series with per-case telemetry attached, and it shows why: p95 latency of 311.6 seconds per case, driven by an average of 5,784 hidden reasoning tokens against roughly 73 tokens of actual answer. Two more frontier models, zero movement on the value ceiling — and the first hard evidence that accuracy and latency trade off independently of it.

Methodology

Ground truth by construction; every verdict, category, and exception value independently reconciled by a separate Python recompute before publication. Nine models, three runs each, deterministic scoring on detection, false-break rate, category attribution, field identification, value extraction, substitution validity, and escalation. A pre-publication gate re-runs the recompute, the scorer contract tests, and per-model error checks; all nine passed at 250 cases × 3 runs, eight with zero errors and Qwen 3.8-Max with 12 (well within the gate's error-rate envelope).

How to cite

BibTeX: @dataset{aal_d_006_2026, title={AAL-D-006: Collateral Eligibility & Substitution Dataset}, author={{AI Alpha Labs}}, year={2026}, month={aug}, version={1.1}, publisher={AI Alpha Labs}, url={https://aialphalabs.ai/research/AAL-D-006}, note={250 cases, 162 ineligible across 12 eligibility categories and 88 traps; nine models; 6,750 scored observations}}. APA style: AI Alpha Labs. (2026). AAL-D-006: Collateral Eligibility & Substitution Dataset (Version 1.1). Retrieved from https://aialphalabs.ai/research/AAL-D-006.

View materials on GitHubBack to portalView full series →
Related