AAL-D-003 — Equity-Options Confirmation Exceptions
AAL-D-003 tests trade-confirmation exception identification on equity options: single-name and index underlyings across Bloomberg-confirmed OTC, CME, and Eurex venues, with multi-leg spreads, dividend schedules, greeks, and corporate-action adjustments. 162 of 250 cases carry an injected exception across 19 categories; 88 are clean traps. Every numeric ground truth is produced by a deterministic pricing engine (Boyle-averaged binomial trees) and independently verified by a recompute checker written from the specification alone. Four models — Claude Sonnet 5, Claude Sonnet 4.6, Gemini 2.5 Pro, and GPT-4o — were evaluated three times per case under two prompt versions, giving 6,000 scored observations and a controlled measurement of how much of model failure is prompt underspecification versus genuine capability limits.
Construction and QA
Cases are generated deterministically from a locked specification, then pass seven QA gates: in-generator invariants, an independent arithmetic recompute (750/750 case instances re-priced from the spec with zero failures), distribution assertions against the planning table, and an adversarial AI subject-matter review that caught two classes of defect the mechanical gates could not — economically impossible contracts in the clean pool and injected greek errors below detection tolerance. Dual-exception cases assign the primary label by a codified risk rule (exposure band plus type-criticality floor), and all relabeling decisions are documented in the audit trail.
Scoring
A deterministic scorer grades detection, category, field, numeric values, exposure (±$1,000), dual-exception recall, escalation, greek identification, and leg attribution. A post-run scoring audit found and fixed two scorer artifacts before publication — exact-string field matching that penalized equivalent field paths, and greek credit restricted to the primary slot — and all results were rescored from stored predictions with the pre-fix outputs archived.
Headline results
Detection is saturated: all four models scored 3,000/3,000 on exception detection with zero false positives on the 88 clean traps. The discriminating metrics sit downstream. Under the v1.0 prompt, no model exceeded 48.8% on exposure calculation. Under v1.1 — which specifies that term breaks carry re-valuation impact and that category means root cause, not symptom — exposure roughly doubled (Sonnet 4.6 63.0%, Sonnet 5 62.8%, Gemini 59.9%, GPT-4o 39.1%) and null exposure answers collapsed (Gemini 44% to 0%). Sonnet 5 reached 94.4% category accuracy. The corporate-action category became a capability ladder: told explicitly to label root cause, Sonnet 5 reaches 70%, Gemini 23%, Sonnet 4.6 13%, GPT-4o 0%.
What persists
The residual 37–40% exposure failure rate for the top three models survives full prompt specification: multiplier slips, per-contract versus total confusion, and re-valuation errors. This replicates the AAL-D-001 and AAL-D-002 finding at higher stakes — extraction is solved, instructions recover recognition, arithmetic does not follow. The dataset and ground truth remain private to prevent training-data contamination; the methodology, scorecards, and audit trail are documented in the results report.
Sample case preview
Case AAL-D-003-089 · Equity Index Option · Multi-Leg Spread · Greek Discrepancy. Four-leg iron condor on SPX, CME venue, expiry Sep 2026. Legs: long 1x 4500 call, short 1x 4600 call, short 1x 4400 put, long 1x 4300 put. Confirmed vega: 45.2; internal vega: 42.1 (root cause: internal used 30-day vol instead of full-term vol surface). Confirmed theta: -18.5; internal theta: -18.5 (match). Ground truth flags one exception: DIS-GREEK, field vega, with exposure $3,100.00. Scoring criteria: category exact; field exact; vega within ±0.5; exposure within $1,000.00; no false positives.
How to cite
BibTeX: @dataset{aal_d_003_2026, title={AAL-D-003: Equity-Options Confirmation Exceptions Dataset}, author={{AI Alpha Labs}}, year={2026}, month={jul}, version={1.0.2}, publisher={AI Alpha Labs}, url={https://aialphalabs.ai/research/AAL-D-003}, note={250 cases across 19 exception categories with controlled prompt ablation across 4 models}}. APA style: AI Alpha Labs. (2026). AAL-D-003: Equity-Options Confirmation Exceptions Dataset (Version 1.0.2). Retrieved from https://aialphalabs.ai/research/AAL-D-003. MLA style: AI Alpha Labs. AAL-D-003: Equity-Options Confirmation Exceptions Dataset, version 1.0.2, July 2026, aialphalabs.ai/research/AAL-D-003.
