Financial arithmetic weakness is structural, not workflow-specific
In AAL-D-001, GPT-4o achieved 68.2% exposure accuracy. The task was secondary: detection and classification were the primary graded dimensions. AAL-D-002 was designed to make arithmetic the primary deliverable. Amount accuracy on AAL-D-002 is 76.1% — confirming that the weakness is not specific to FX rate arithmetic in trade confirmations, but is a structural limitation in financial multi-step calculation across workflows. Models can reliably identify what is wrong; they cannot reliably compute what the correct number should be.
What changes across datasets and what does not
Detection accuracy improves slightly from 99.2% (AAL-D-001) to 99.9% (AAL-D-002). Amount accuracy improves modestly from 68.2% to 76.1%, likely because AAL-D-002 amounts involve simpler integer arithmetic versus FX rate pip calculations. The structural gap — models detect at ~99% and calculate at ~70–76% — is unchanged across both datasets.
Operational implication
For any capital markets operations workflow where a correct dollar figure is the output, model-generated amounts require validation against system-of-record calculations. The model is a reliable triage tool and an unreliable calculator. Deployment architectures should route detected exceptions to deterministic calculation engines for amount verification rather than relying on the model's arithmetic.
Path to confirmation across models
This finding is currently confirmed for GPT-4o across two datasets. Gemini and Claude evaluations on AAL-D-002 are planned as AAL-RS-008 and AAL-RS-009. If the pattern holds across model families — as the AAL-D-001 data suggests it will — this advances to a confirmed cross-model structural finding with direct implications for production deployment design.
