Writing

Research notes and findings — in plain language.

Longer-form write-ups of what we found, how we found it, and what it means for capital markets operations. Each post links to the underlying research entries that support it.

Can AI Read Trade Confirmations? We Tested It.

We ran GPT-4o, Gemini 2.5 Pro, Gemini 2.5 Flash, and Claude Sonnet on the same 250-case trade confirmation benchmark. Here's what we found — and what we didn't expect.

benchmarkevaluationmethodology

LLM Benchmark Contamination: How a Quota Error Corrupted 29% of Our Gemini Eval

218 contaminated runs, null predictions, and why audit gates matter more than perfect pipelines.

auditmethodologyevaluationgeminicontamination

LLM Arithmetic Failures in Capital Markets: How We Caught and Fixed the Anti-Pattern

An automated review caught our dispute endpoint doing the exact thing our benchmarks warn against — LLM arithmetic. Here's the fix, and the replay methodology that measured exactly what changed.

methodologyarchitecturebenchmark

We Told Four Frontier Models Exactly What to Do. Half the Failure Vanished. The Other Half Is Arithmetic.

AAL-D-003 put equity-options confirmations in front of Claude Sonnet 5, Sonnet 4.6, Gemini 2.5 Pro, and GPT-4o — twice, with one prompt change in between. The controlled result: specification recovers recognition, but the arithmetic gap survives explicit instruction.

benchmarkevaluationmethodologyoptions

Reasoning Buys Computation, Not Convention: What Settlement-Fail Traps Reveal

AAL-D-004 put settlement failures in front of a thinking model and three non-thinking ones. Detection saturated at 100%. The discriminator was the false alarm — and the only model that reasons its way out does so on the math traps, not the convention traps.

benchmarkevaluationmethodologysettlementreasoning

The Scarce Resource in Capital Markets AI Isn't Accuracy. It's Reproducibility.

Most AI accuracy numbers in capital markets procurement can't be reproduced — vague model versions, undisclosed datasets, proprietary rubrics, single runs. We built 1,000 cases across four datasets for about $150–200. If rigorous evidence is that cheap, its absence from a vendor claim is a choice. Here's a 10-question scorecard to catch it.

methodologyreproducibilityprocurementbenchmark
Get the next finding by email.

The AI Alpha Brief delivers new benchmark results, model limitations and implementation implications for capital-markets teams.

Subscribe to the Brief →