Independent AI evaluation · Capital markets

We evaluate AI for capital markets.

AI Alpha Labs tests frontier models on real trade-operations workflows — scored by a deterministic engine, published with full methodology, and reproducible from the materials we release.

1,500
Benchmark cases across AAL-D-001 through D-006
9
Frontier models on AAL-D-006 — GPT-5.6 Sol, Opus 5, Sonnet 5, Gemini 3.1 Pro, Grok 4.5, DeepSeek V4, Kimi K3, Muse Spark 1.2, Qwen 3.8-Max
99.9%
GPT-4o margin call dispute detection accuracy
4
Live API endpoints — validate, validate-raw, batch, disputes
Why AI Alpha Labs

Not a model vendor. An independent evaluator.

01

Independent

We don't build the models we score. No commercial incentive to inflate a result — the only product is the evidence.

02

Reproducible

Every benchmark ships with its dataset, prompts, and a deterministic scorer. Re-run it and get the same number.

03

Transparent

We publish confidence intervals, variance, and the cases models fail — not just a headline accuracy figure.

04

Financial-services focused

Benchmarks built on real trade-operations workflows — confirmations, exceptions, settlement — not academic tasks.

Latest Research

A growing body of evidence.

All research →Series Roadmap →
AAL-R-2026-001Published

Trade Confirmation Exception Identification

The flagship study: can frontier models detect, classify, and quantify settlement exceptions across seven asset classes, scored deterministically?

v1.1 · Jul 2026Read →
AAL-D-001Dataset

Benchmark Overview & Dataset

250 validated cases pairing counterparty confirmations against internal records, with ground truth and per-case scoring criteria.

250 cases · v1.0Explore →
AAL-RS-001Result

Benchmark Results — Four Models, Three Labs

Four published evaluations across GPT-4o, Gemini 2.5 Pro, Gemini 2.5 Flash, and Claude Sonnet. Detection range 97.9–99.6%. Failures documented, not hidden.

4 models · Jul 2026View →
AAL-D-002Dataset

Margin Call Dispute Detection

250 cases across 17 dispute categories — DIS-PRICE, DIS-HAIRCUT, DIS-THRESH, DIS-SIMM and more. The first public benchmark for LLM margin call dispute resolution.

250 cases · v1.0Explore →
AAL-RS-007Published

Margin Call Disputes — GPT-4o vs Claude Sonnet

GPT-4o achieves 99.9% dispute detection and 76.1% amount accuracy. The 23.8pp arithmetic gap replicates the core finding from AAL-D-001 across a different workflow.

2 models · Jul 2026View →
AAL-D-003Published

Equity-Options Confirmation Exceptions

250 equity-options cases across 19 exception categories. Detection saturates at 100% across all four models. The discriminating metric is exposure calculation — no model exceeded 48.8% under v1.0 prompting.

4 models · 6,000 obs · Jul 2026Explore →
AAL-D-004Published

Settlement Fail Root Cause Identification

250 settlement cases, 88 of them trap cleans. All four models catch 100% of real exceptions — false-positive rate is the discriminator: 15.2% (Sonnet 5, thinking) vs 38.6% (Sonnet 4.6, GPT-4o) vs 49.6% (Gemini 2.5 Pro).

4 models · 3,000 obs · Jul 2026Explore →
AAL-D-006Published

Collateral Eligibility & Substitution

250 cases across 12 eligibility categories, on nine frontier models. Detection is solved (93–100%); no model values the exposure above 58%; and escalation splits the field from 5.6% to 98.8%.

9 models · Aug 2026Explore →
AAL-D-005Published

P&L Reconciliation Break Attribution

250 cases — 162 breaks across 15 root-cause categories, 88 explained-move traps — across six models. Detection is solved (94.8–98.9%); value extraction splits the field from 55.4% to 11.9%, and “reasoning model” proves to be a label that doesn't predict performance.

6 models · 4,500 obs · Aug 2026Explore →
AAL-RS-010Published

The Reproducibility Crisis in Capital Markets AI Benchmarking

Most vendor AI accuracy claims can't be reproduced. This note contrasts that with the AAL corpus — 1,000 synthetic cases, 4 datasets, deterministic scoring, ~$150–200 — and ends with a 10-question Reproducibility Scorecard you can send any AI vendor.

Publication · Jul 2026Read →
The AI Alpha Brief

Benchmark findings, documented model failures and practical implications for capital-markets teams — delivered weekly.

Subscribe to the Brief →
AAL Benchmark Series · 1,500 cases · six datasets

AI can spot the problem. It can't always calculate the answer.

Each chart shows two bars per model. Navy = detection accuracy — did the model correctly identify whether a problem exists? Gold = the discriminating metric — calculation accuracy (D-001–003), trap-clean accuracy (D-004), or value accuracy (D-005–006). The gap between them is the finding.
AAL-D-001 · Trade confirmation exception identification
Can the model catch a discrepancy between a counterparty confirmation and an internal record — and quantify the exposure?
AAL-D-002 · Margin call dispute detection
Can the model identify a disputed margin call — and compute the correct call amount after applying threshold, haircut, and MTA rules?
AAL-D-003 · Equity-options confirmation exceptions · v1.1 prompt
Can the model detect an exception in an equity options confirmation — and calculate the correct exposure, including greeks and multi-leg structures?
AAL-D-004 · Settlement fail root cause identification
Can the model catch a real settlement failure — without crying wolf on trap cleans (settlements that look wrong but aren't)? Gold bar = trap-clean accuracy (100 − false-positive rate).
AAL-D-005 · P&L reconciliation break attribution
Can the model tell a real P&L break from a fully-explained move — and pin the actual dollar figure? Gold bar = value accuracy. Note the spread: two 'reasoning' models (Sonnet 5, KIMI K3) sit 12 points apart, and a third (DeepSeek) lands below a non-reasoning model — the label doesn't predict the number.
AAL-D-006 · Collateral eligibility & substitution
Can the model decide whether collateral is eligible, value the ineligible exposure, and propose a valid substitute? Gold bar = value accuracy, on the seven newest frontier models. No model clears 58% on the number — and no single model wins across the board: Gemini leads value, Opus 5 escalation (98.8%), Grok 4.5 substitution (98.2%), yet each hits the same wall.
The finding:Detection is a solved problem — frontier models identify exceptions at 93–100% across all six datasets. Judgment is not. No model has exceeded 76% on financial calculation accuracy; on D-005 value extraction spans 55.4% down to 11.9%, and on D-006 collateral eligibility no model clears 58% — even “reasoning” models scatter across the range, and no single model wins across the board. This is why the AAL API routes classification to an LLM and all arithmetic to a deterministic Python engine.
All evaluations at temperature 0, 3 runs per case, Wilson 95% CI. Full methodology at aialphalabs.ai/research
Full resultsAll research →
AAL-D-001 · Trade Confirmation Exception Identification250 cases · 7 asset classes →
#ModelDetectionExposure acc.False pos.Status
01Gemini 2.5 Prov1.2 prompt99.6~76%0.0%Published
02GPT-4ov1.2 prompt99.268.2%0.0%Published
03Gemini 2.5 Flashv1.2 prompt99.2~76%0.0%Published
04Claude Sonnet 4.6v1.2 prompt98.863.5%0.0%Published
AAL-D-002 · Margin Call Dispute Detection250 cases · 17 dispute categories →
#ModelDetectionAmount acc.False pos.Status
01GPT-4ogpt-4o-2024-08-0699.9%76.1%0.0%Published
02Claude Haikuproduction API81.6%~40%0.8%Published
03Gemini 2.5 Proaudit-confirmed65.6%44.5%9.0%Published
AAL-D-003 · Equity-Options Confirmation Exceptions250 cases · 19 categories · 6,000 obs →
#ModelDetectionExposure acc. v1.1Category acc. v1.1Status
01Claude Sonnet 5claude-sonnet-5100%62.8%94.4%Published
02Claude Sonnet 4.6claude-sonnet-4-6100%63.0%Published
03Gemini 2.5 Progemini-2.5-pro100%59.9%Published
04GPT-4ogpt-4o-2024-08-06100%39.1%Published
AAL-D-004 · Settlement Fail Root Cause Identification250 cases · 88 trap cleans · 3,000 obs →
#ModelDetectionFalse-pos. rateCategory acc.Status
01Claude Sonnet 5thinking · claude-sonnet-5100%15.2%Published
02Claude Sonnet 4.6claude-sonnet-4-6100%38.6%86.1%Published
03GPT-4ogpt-4o-2024-08-06100%38.6%Published
04Gemini 2.5 Progemini-2.5-pro100%49.6%Published
AAL-D-005 · P&L Reconciliation Break Attribution250 cases · 162 breaks / 88 traps · 4,500 obs →
#ModelDetectionValue acc.False-breakStatus
01Claude Sonnet 5thinking · claude-sonnet-598.4%55.4%Published
02KIMI K3thinking · kimi-k398.0%43.4%0.4%Published
03Claude Sonnet 4.6claude-sonnet-4-698.4%39.4%Published
04Gemini 2.5 Progemini-2.5-pro98.1%30.3%Published
05DeepSeek V3.2thinking · deepseek-v3.294.8%26.9%Published
06GPT-4ogpt-4o-2024-08-0698.9%11.9%0.0%Published
AAL-D-006 · Collateral Eligibility & Substitution250 cases · 162 ineligible / 88 traps · 9 models →
#ModelDetectionValue acc.EscalationStatus
01Gemini 3.1 Progemini-3.1-pro-preview100%57.8%86.0%Published
02Qwen 3.8-Maxdashscope · qwen3.8-max100%39.2%51.9%Published
03GPT-5.6 Solgpt-5.6-sol100%38.3%79.8%Published
04Claude Sonnet 5thinking · claude-sonnet-599.7%35.5%98.4%Published
05Kimi K3thinking · kimi-k399.9%31.6%69.8%Published
06Muse Spark 1.2meta/muse-spark-1.2100%31.1%83.5%Published
07DeepSeek V4 Prothinking · deepseek-v4-pro93.1%24.0%5.6%Published
08Claude Opus 5thinking · claude-opus-595.7%22.5%98.8%Published
09Grok 4.5x-ai/grok-4.5100%15.4%54.1%Published

All evaluations use deterministic scoring engines with per-case tolerances. Wilson 95% confidence intervals on all accuracy figures. Full methodology, prompts, and rubrics published with each dataset.

Evaluation principle
We don't benchmark models. We benchmark models on the work your desk actually does.

General leaderboards measure academic capability. AAL-D-001 measures trade-confirmation exception handling under production conditions.

Philosophy

Eight principles.

01

Typography is the brand.

Information presented with precision builds more trust than decoration. We let the work speak.

02

Data before decoration.

Every element earns its place by carrying information. Visual complexity that adds no meaning is removed.

03

Whitespace is confidence.

Density signals anxiety. Clarity signals command. We optimize for the reader, not the page.

04

Motion is subtle.

Animation that calls attention to itself is a distraction. Interfaces move only when movement carries meaning.

05

Every page is printable.

If content can't stand without interactive chrome, we reconsider the content.

06

Components earn their place.

We don't add UI because it looks standard. We add it because the reader needs it.

07

Consistency builds trust.

Predictable patterns lower cognitive load. Every result is read the same way as the last.

08

Simplicity scales.

Simple principles outlast clever systems. We optimize for reproducibility and extension.

Custom benchmarks and private briefings for institutional operations and risk teams.

Contact for access