Independent AI evaluation for capital markets operations.
AI Alpha Labs was founded to bring independent, reproducible benchmarking to a market where every AI vendor publishes its own accuracy numbers and nobody audits the methodology.
Capital markets operations is a consequential domain. A missed margin call dispute, a misbooked trade confirmation, a settlement failure that wasn't flagged — these are not abstract errors. The AI being evaluated for these workflows deserves evaluation by someone with no stake in the result.
AAL builds benchmark datasets from real capital markets operations workflows — trade confirmation matching, margin call dispute resolution, settlement exception identification. Each dataset contains hundreds of cases with pre-constructed ground truth, scored by a deterministic engine rather than model judgment.
We run frontier AI models — GPT-4o, Gemini 2.5 Pro, Claude Sonnet — against these datasets and publish the results with full methodology. Every evaluation uses multiple runs, Wilson confidence intervals, and a contamination audit gate before any result is published.
The research is free and public. The commercial products — AAL Certified for vendors, Buy-side Audits for firms — apply the same methodology to external systems under contract.
Detection is solved. Arithmetic is not.
Across 1,500 benchmark cases and six datasets, the pattern is consistent: AI detects exceptions at 93–100% accuracy. The same models compute correct dollar amounts at 12–76% accuracy. The gap between those two numbers — 20 to 85 percentage points — is persistent across different asset classes, different workflows, and different models.
This finding drives the architecture of the AAL API: language models handle classification and extraction; a deterministic Python engine handles all arithmetic. The benchmark is the evidence for the architecture. See the published results →
How we work.
Independence
AI Alpha Labs does not build the models it evaluates. It does not sell AI products or take commissions from model vendors. The only commercial relationship is with the firms and vendors who commission benchmarks and audits. Every result is published with full methodology so anyone can verify it.
Deterministic scoring
Every benchmark result is scored by a deterministic Python engine — not by a language model judging another language model. Ground truth is constructed before any model sees a case. Scoring criteria are fixed before evaluation begins. Nothing is changed after results are known.
Reproducibility
Every published result includes the dataset schema, scoring methodology, prompt versions, model identifiers, and evaluation dates. If a result cannot be reproduced from the published materials, it should not be trusted. AAL holds itself to the same standard.
Honest failure reporting
Model failures are documented, not hidden. When an eval pipeline produces contaminated results, the contamination is disclosed and the data is purged before publication. When a model underperforms, the result is published. The benchmark is only credible if it can produce a failing grade.
Six datasets. 1,500 cases. All public.
Questions about the methodology or the research? We publish everything — but we're also happy to discuss it directly.
joe@aialphalabs.ai