Research notes and findings — in plain language.
Longer-form write-ups of what we found, how we found it, and what it means for capital markets operations. Each post links to the underlying research entries that support it.
Can AI Read Trade Confirmations? We Tested It.
We ran GPT-4o, Gemini 2.5 Pro, Gemini 2.5 Flash, and Claude Sonnet on the same 250-case trade confirmation benchmark. Here's what we found — and what we didn't expect.
LLM Benchmark Contamination: How a Quota Error Corrupted 29% of Our Gemini Eval
218 contaminated runs, null predictions, and why audit gates matter more than perfect pipelines.
LLM Arithmetic Failures in Capital Markets: How We Caught and Fixed the Anti-Pattern
An automated review caught our dispute endpoint doing the exact thing our benchmarks warn against — LLM arithmetic. Here's the fix, and the replay methodology that measured exactly what changed.
We Told Four Frontier Models Exactly What to Do. Half the Failure Vanished. The Other Half Is Arithmetic.
AAL-D-003 put equity-options confirmations in front of Claude Sonnet 5, Sonnet 4.6, Gemini 2.5 Pro, and GPT-4o — twice, with one prompt change in between. The controlled result: specification recovers recognition, but the arithmetic gap survives explicit instruction.
Reasoning Buys Computation, Not Convention: What Settlement-Fail Traps Reveal
AAL-D-004 put settlement failures in front of a thinking model and three non-thinking ones. Detection saturated at 100%. The discriminator was the false alarm — and the only model that reasons its way out does so on the math traps, not the convention traps.
The Scarce Resource in Capital Markets AI Isn't Accuracy. It's Reproducibility.
Most AI accuracy numbers in capital markets procurement can't be reproduced — vague model versions, undisclosed datasets, proprietary rubrics, single runs. We built 1,000 cases across four datasets for about $150–200. If rigorous evidence is that cheap, its absence from a vendor claim is a choice. Here's a 10-question scorecard to catch it.
The AI Alpha Brief delivers new benchmark results, model limitations and implementation implications for capital-markets teams.
Subscribe to the Brief →