Every benchmark, evaluation, and finding — live.
The complete AI Alpha Labs corpus. Each entry carries its status, its evidence level, and a link to the materials that make it reproducible. Nothing is published until it is defensible.
Trade Confirmation Exception Identification
A production benchmark for whether frontier models can detect, classify, and quantify settlement exceptions across seven asset classes.
Trade Confirmation Exception Identification — Dataset
250 validated cases pairing counterparty confirmations against internal records across cash and derivative products, with ground truth and per-case scoring criteria.
Margin Call Dispute Detection — Dataset
250 validated cases covering margin call dispute detection, classification, and correct-amount calculation across OTC derivatives, exchange-cleared products, and securities lending.
AAL-D-003 — Equity-Options Confirmation Exceptions
250 equity-options confirmation cases with greeks, multi-leg spreads, and dual exceptions — plus a controlled prompt ablation across four frontier models.
Settlement Fail Root Cause Identification — Equity Options
250 equity-options settlement failure cases with trap cleans that separate thinking from non-thinking models — the first benchmark where false-positive rate, not detection, discriminates the field.
P&L Reconciliation Break Attribution
250 daily P&L reconciliation cases — 162 breaks across 15 root-cause categories, 88 explained-move traps — across six frontier models including DeepSeek V3.2 and KIMI K3. Detection is solved; quantification splits the field from 55% to 12%; and 'reasoning model' proves to be a label that does not predict performance.
Collateral Eligibility & Substitution
250 collateral cases — 162 ineligible across 12 eligibility categories, 88 traps — across nine frontier models. Detection is solved (93–100%); no model values the exposure above 58%; escalation splits the field from 5.6% to 98.8%. A companion cut tests whether a 30B open-weight model closes the gap — it doesn't.
SIMM / UMR Initial Margin Dispute Detection
250 bilateral initial-margin reconciliation cases — 162 disputes across 9 IM-DIS categories, 88 within-tolerance traps — under ISDA SIMM, across eight frontier models. Detection is not merely solved but stable: 100.0% on every model, 0.0% false flags, and zero run-to-run contradictions across all 250 cases. Naming the cause of the dispute collapses to 45.7–63.4%, and run-to-run value instability spans 5.6% to 40.1%.
GPT-4o on AAL-D-001
First published evaluation: 97.9% first-pass detection across all 250 cases, with the single failure case documented rather than hidden.
Gemini 2.5 Pro on AAL-D-001
99.6% detection, 100% dual-exception recall, 0 errors across 750 scored observations — and a persistent exposure arithmetic weakness that survives prompt improvement.
Gemini 2.5 Flash on AAL-D-001
99.2% detection, 100% dual-exception recall, zero errors across 750 observations — statistically indistinguishable from Pro on every metric in v1.1.
Prompt Sensitivity Analysis: v1.0 vs v1.1 on AAL-D-001
A 19-point escalation gap between Pro and Flash vanishes with one paragraph. Exposure arithmetic stays broken. Benchmark iteration separates specification gaps from genuine model limitations.
Claude Sonnet 4.6 on AAL-D-001
98.8% detection, 0.0% false positive rate, 0 errors across 750 observations — with a lower exposure accuracy (63.5%) and dual-exception recall (87.4%) than Gemini, and a novel EXC-STAT false positive pattern requiring a v1.2 prompt fix.
GPT-4o on AAL-D-001 (v1.2 multi-run)
99.2% detection, 100% dual-exception recall, 0 errors across 750 observations — closing the comparison table with full sub-metrics and confirming GPT-4o matches Gemini on all dimensions except exposure arithmetic.
GPT-4o on AAL-D-002
99.9% detection, 94.9% category accuracy, 0.0% false positive rate across 750 observations — and amount accuracy of 76.1%, confirming the arithmetic weakness from AAL-D-001 survives into a workflow where correct calculation is the primary deliverable.
The Economics of LLM Deployment in Capital Markets Operations
The total API cost to build the AAL benchmark corpus — 750 cases, 5 frontier models, 3 datasets at the time of writing (since grown to 1,000 cases across four datasets) — was approximately $150–200. Commercial products built on that evidence base operate at 98–99% gross margins. The margin profile is structurally determined by an architecture that routes classification to LLMs and arithmetic to deterministic Python.
The Reproducibility Crisis in Capital Markets AI Benchmarking
Most AI accuracy figures cited in capital markets procurement are neither reproducible nor independently validated. This note contrasts that norm with the AAL protocol — 1,000 synthetic-by-construction cases across four datasets, up to four frontier models at three runs each, all arithmetic in deterministic Python, built for ~$150–200 — and provides a vendor due-diligence checklist and a 10-question Reproducibility Scorecard for procurement teams. Reproducibility, not accuracy, is the scarce resource.
Detection accuracy clusters near ceiling on objective exceptions
Across four independent model evaluations across three labs, first-pass detection on objective exceptions sits at 97.9–99.6% — the pattern is now confirmed across models, architectures, and labs.
Residual errors concentrate in judgment-heavy booking exceptions
Three independent frontier models miss the same judgment-heavy derivative cases while clearing 99%+ of everything else — residual risk lives in entity- and structure-level reasoning.
Models detect exceptions far more reliably than they quantify them
Detection sits at ~99.5% while exposure-arithmetic accuracy plateaus at ~75% across both models and both prompt versions — a 25-point gap that does not close with explicit formula guidance.
Benchmark specification gaps can masquerade as model capability differences
A 19-point escalation gap between Pro and Flash in v1.0 vanished entirely in v1.1 after adding one paragraph of policy guidance — a case study in how underspecified benchmarks produce misleading model comparisons.
Dual-exception recall and exposure accuracy vary meaningfully across model families
GPT-4o and Gemini models achieve 100% dual-exception recall; Claude Sonnet achieves 87.4%. Exposure accuracy splits by lab: Gemini ~76%, GPT-4o 68.2%, Claude 63.5% — a confirmed model-family pattern across five evaluations.
Financial arithmetic weakness is structural, not workflow-specific
Amount accuracy at 76.1% on AAL-D-002 — a workflow where correct calculation is the primary deliverable — confirms that the arithmetic limitation first observed in AAL-D-001 is structural, not an artifact of the trade confirmation task.
Claude Sonnet 4.6 on AAL-D-002
99.3% detection, 92.2% category accuracy, 0.0% false positive rate across 750 observations — and amount accuracy of 80.4%, edging GPT-4o on every arithmetic metric while confirming the structural ceiling holds across both frontier models.
Arithmetic ceiling confirmed across two frontier models on AAL-D-002
GPT-4o at 76.1% and Claude Sonnet at 80.4% on correct call amounts — against ≥99.3% detection and ≥92% category accuracy — confirm that the detection/arithmetic gap is structural across both frontier model families on the margin call workflow.
A 30B open-weight model matches frontier detection — and hits the same value wall
Meta's Muse Glimmer (30B, open-weight, single-GPU local deployment) posts 95.7% detection on AAL-D-006 — beating a closed frontier flagship — and 78.6% substitution validity, beating two flagships. Value accuracy: 16.3%, the same ceiling every model in the corpus hits regardless of scale.
