Research Portal

Every benchmark, evaluation, and finding — live.

The complete AI Alpha Labs corpus. Each entry carries its status, its evidence level, and a link to the materials that make it reproducible. Nothing is published until it is defensible.

Methodology →Series Roadmap →Reproducibility Scorecard →
AAL-R-2026-001Publications

Trade Confirmation Exception Identification

A production benchmark for whether frontier models can detect, classify, and quantify settlement exceptions across seven asset classes.

Published
AAL-D-001Benchmarks

Trade Confirmation Exception Identification — Dataset

250 validated cases pairing counterparty confirmations against internal records across cash and derivative products, with ground truth and per-case scoring criteria.

Active · v1.0
AAL-D-002Benchmarks

Margin Call Dispute Detection — Dataset

250 validated cases covering margin call dispute detection, classification, and correct-amount calculation across OTC derivatives, exchange-cleared products, and securities lending.

Active · v1.0
AAL-D-003Benchmarks

AAL-D-003 — Equity-Options Confirmation Exceptions

250 equity-options confirmation cases with greeks, multi-leg spreads, and dual exceptions — plus a controlled prompt ablation across four frontier models.

Published
AAL-D-004Benchmarks

Settlement Fail Root Cause Identification — Equity Options

250 equity-options settlement failure cases with trap cleans that separate thinking from non-thinking models — the first benchmark where false-positive rate, not detection, discriminates the field.

Published
AAL-D-005Benchmarks

P&L Reconciliation Break Attribution

250 daily P&L reconciliation cases — 162 breaks across 15 root-cause categories, 88 explained-move traps — across six frontier models including DeepSeek V3.2 and KIMI K3. Detection is solved; quantification splits the field from 55% to 12%; and 'reasoning model' proves to be a label that does not predict performance.

Published
AAL-D-006Benchmarks

Collateral Eligibility & Substitution

250 collateral cases — 162 ineligible across 12 eligibility categories, 88 traps — across nine frontier models. Detection is solved (93–100%); no model values the exposure above 58%; escalation splits the field from 5.6% to 98.8%. A companion cut tests whether a 30B open-weight model closes the gap — it doesn't.

Published
AAL-D-007Benchmarks

SIMM / UMR Initial Margin Dispute Detection

250 bilateral initial-margin reconciliation cases — 162 disputes across 9 IM-DIS categories, 88 within-tolerance traps — under ISDA SIMM, across eight frontier models. Detection is not merely solved but stable: 100.0% on every model, 0.0% false flags, and zero run-to-run contradictions across all 250 cases. Naming the cause of the dispute collapses to 45.7–63.4%, and run-to-run value instability spans 5.6% to 40.1%.

Published
AAL-RS-001Evaluations

GPT-4o on AAL-D-001

First published evaluation: 97.9% first-pass detection across all 250 cases, with the single failure case documented rather than hidden.

Published
AAL-RS-002Evaluations

Gemini 2.5 Pro on AAL-D-001

99.6% detection, 100% dual-exception recall, 0 errors across 750 scored observations — and a persistent exposure arithmetic weakness that survives prompt improvement.

Published
AAL-RS-003Evaluations

Gemini 2.5 Flash on AAL-D-001

99.2% detection, 100% dual-exception recall, zero errors across 750 observations — statistically indistinguishable from Pro on every metric in v1.1.

Published
AAL-RS-004Evaluations

Prompt Sensitivity Analysis: v1.0 vs v1.1 on AAL-D-001

A 19-point escalation gap between Pro and Flash vanishes with one paragraph. Exposure arithmetic stays broken. Benchmark iteration separates specification gaps from genuine model limitations.

Published
AAL-RS-005Evaluations

Claude Sonnet 4.6 on AAL-D-001

98.8% detection, 0.0% false positive rate, 0 errors across 750 observations — with a lower exposure accuracy (63.5%) and dual-exception recall (87.4%) than Gemini, and a novel EXC-STAT false positive pattern requiring a v1.2 prompt fix.

Published
AAL-RS-006Evaluations

GPT-4o on AAL-D-001 (v1.2 multi-run)

99.2% detection, 100% dual-exception recall, 0 errors across 750 observations — closing the comparison table with full sub-metrics and confirming GPT-4o matches Gemini on all dimensions except exposure arithmetic.

Published
AAL-RS-007Evaluations

GPT-4o on AAL-D-002

99.9% detection, 94.9% category accuracy, 0.0% false positive rate across 750 observations — and amount accuracy of 76.1%, confirming the arithmetic weakness from AAL-D-001 survives into a workflow where correct calculation is the primary deliverable.

Published
AAL-RS-009Publications

The Economics of LLM Deployment in Capital Markets Operations

The total API cost to build the AAL benchmark corpus — 750 cases, 5 frontier models, 3 datasets at the time of writing (since grown to 1,000 cases across four datasets) — was approximately $150–200. Commercial products built on that evidence base operate at 98–99% gross margins. The margin profile is structurally determined by an architecture that routes classification to LLMs and arithmetic to deterministic Python.

Published
AAL-RS-010Publications

The Reproducibility Crisis in Capital Markets AI Benchmarking

Most AI accuracy figures cited in capital markets procurement are neither reproducible nor independently validated. This note contrasts that norm with the AAL protocol — 1,000 synthetic-by-construction cases across four datasets, up to four frontier models at three runs each, all arithmetic in deterministic Python, built for ~$150–200 — and provides a vendor due-diligence checklist and a 10-question Reproducibility Scorecard for procurement teams. Reproducibility, not accuracy, is the scarce resource.

Published
AAL-F-001Findings

Detection accuracy clusters near ceiling on objective exceptions

Across four independent model evaluations across three labs, first-pass detection on objective exceptions sits at 97.9–99.6% — the pattern is now confirmed across models, architectures, and labs.

Confirmed
AAL-F-002Findings

Residual errors concentrate in judgment-heavy booking exceptions

Three independent frontier models miss the same judgment-heavy derivative cases while clearing 99%+ of everything else — residual risk lives in entity- and structure-level reasoning.

Confirmed
AAL-F-003Findings

Models detect exceptions far more reliably than they quantify them

Detection sits at ~99.5% while exposure-arithmetic accuracy plateaus at ~75% across both models and both prompt versions — a 25-point gap that does not close with explicit formula guidance.

Confirmed
AAL-F-004Findings

Benchmark specification gaps can masquerade as model capability differences

A 19-point escalation gap between Pro and Flash in v1.0 vanished entirely in v1.1 after adding one paragraph of policy guidance — a case study in how underspecified benchmarks produce misleading model comparisons.

Confirmed
AAL-F-005Findings

Dual-exception recall and exposure accuracy vary meaningfully across model families

GPT-4o and Gemini models achieve 100% dual-exception recall; Claude Sonnet achieves 87.4%. Exposure accuracy splits by lab: Gemini ~76%, GPT-4o 68.2%, Claude 63.5% — a confirmed model-family pattern across five evaluations.

Provisional
AAL-F-006Findings

Financial arithmetic weakness is structural, not workflow-specific

Amount accuracy at 76.1% on AAL-D-002 — a workflow where correct calculation is the primary deliverable — confirms that the arithmetic limitation first observed in AAL-D-001 is structural, not an artifact of the trade confirmation task.

Confirmed
AAL-RS-008Evaluations

Claude Sonnet 4.6 on AAL-D-002

99.3% detection, 92.2% category accuracy, 0.0% false positive rate across 750 observations — and amount accuracy of 80.4%, edging GPT-4o on every arithmetic metric while confirming the structural ceiling holds across both frontier models.

Published
AAL-F-007Findings

Arithmetic ceiling confirmed across two frontier models on AAL-D-002

GPT-4o at 76.1% and Claude Sonnet at 80.4% on correct call amounts — against ≥99.3% detection and ≥92% category accuracy — confirm that the detection/arithmetic gap is structural across both frontier model families on the margin call workflow.

Confirmed
AAL-F-008Findings

A 30B open-weight model matches frontier detection — and hits the same value wall

Meta's Muse Glimmer (30B, open-weight, single-GPU local deployment) posts 95.7% detection on AAL-D-006 — beating a closed frontier flagship — and 78.6% substitution validity, beating two flagships. Value accuracy: 16.3%, the same ceiling every model in the corpus hits regardless of scale.

Confirmed