SIMM / UMR Initial Margin Dispute Detection
AAL-D-007 puts bilateral initial-margin reconciliation in front of the frontier field. Under UMR, two counterparties independently compute IM on the same netting set using ISDA SIMM; when the two numbers diverge beyond tolerance, someone has to decide whether a dispute exists and what caused it. The dataset is 250 netting sets — 162 genuine disputes across nine categories at 18 cases each (sensitivity discrepancies, CRIF mapping errors, FX conversion, concentration add-ons, calculation-date staleness, methodology version drift, trade-population gaps, bucket misassignment, netting-set scope) and 88 within-tolerance traps that look like disputes but are not. Version 1.1 corrects a v1.0 design flaw in which the injected error always landed on the counterparty's copy, letting a model score well by copying the firm's figure; v1.1 balances the perturbation 125/125 across sides and removes the prompt hint. All v1.0 results were archived and re-run. Eight models were evaluated — GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.5, DeepSeek V4 Pro, Kimi K3 and Meta's Muse Spark 1.2 — three runs each, deterministic scoring, no model permitted to compute. The finding inverts the usual shape. Detection is saturated: every model at 100.0%, a 0.0% false-flag rate on all 88 traps, and — uniquely in the corpus — a 0.0% run-to-run flip rate, meaning no model ever contradicted itself about whether a dispute exists. The wall has moved to attribution. Category accuracy runs 45.7% (DeepSeek V4 Pro) to 63.4% (Claude Opus 5), and only Opus 5's confidence interval clears 57%. Underneath that sits a second finding: run each case three times and value answers move. The value flip rate spans 5.6% (GPT-5.6 Sol) to 40.1% (DeepSeek V4 Pro) — wider than the value-accuracy spread itself. DeepSeek's 65.4% value accuracy is therefore not the profile of a model wrong a third of the time; it is an average over a partly random process. A benchmark run once cannot distinguish a reliably mediocre model from an unreliable one.
Detection is not just solved — it is stable
All eight models post 100.0% detection accuracy [99.5–100.0] with a 0.0% false-flag rate across all 88 within-tolerance traps. Difference sizing is 100% and escalation is 100% on seven of eight. More striking: across 250 cases and three runs each, no model ever gave a different verdict on whether a dispute exists. Accurate and self-consistent together is the strongest form of this finding anywhere in the corpus.
The wall moved from value to attribution
In D-004 through D-006 the ceiling sat on the dollar figure. Here value accuracy is comparatively strong (65.4–86.6%) while category attribution collapses to 45.7–63.4% — the weakest attribution numbers in the series, and only Claude Opus 5's confidence interval clears 57%. Models can tell you a dispute exists and how large it is; asked why, across nine equally-weighted categories, most of the field is near a coin flip. Opus 5 leads on category (63.4%) and component (82.1%); GPT-5.6 Sol leads on value (86.6%) while posting one of the worst category scores (48.6%). No model leads on both.
A single run cannot tell you what you have
Running each case three times exposes something a pooled accuracy figure hides. Value flip rate — the share of disputed cases where a model gets the IM amount right on some runs and wrong on others — spans 5.6% (GPT-5.6 Sol) to 40.1% (DeepSeek V4 Pro). That range is wider than the value-accuracy range. Two of every five disputed cases produce a different figure from DeepSeek depending on which run you read. Its 65.4% is an average over a partly random process, not a stable capability, and only multi-run methodology surfaces the difference.
Escalation is a property of the task, not the model
Seven of eight models post 100.0% escalation accuracy here; DeepSeek V4 Pro posts 99.2%. The same DeepSeek posted 5.6% escalation on AAL-D-006 — an adjacent middle-office workflow — a 93-point swing on the same model. Escalation reliability cannot be inherited from a vendor scorecard or carried across deployments; it has to be measured on the workflow where the model will run.
v1.1 methodology correction
v1.0 always injected the dispute-causing perturbation into the counterparty's copy, which meant value accuracy could be earned by copying the firm's figure rather than by judgment. The flaw was caught on 19 Aug 2026; all v1.0 results — including a GPT-5.6 Sol run that scored 100/100/100 — were invalidated and archived rather than deleted. v1.1 balances the perturbation 125/125 across firm and counterparty sides and removes the prompt hint indicating which side is authoritative. Every figure published here is from v1.1.
Roster and known limitations
Eight models complete: GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.5, DeepSeek V4 Pro, Kimi K3, Muse Spark 1.2. A ninth, Alibaba's Qwen 3.8-Max, was attempted but not completed; its partial run is excluded rather than reported. Grok 4.5 recorded 22 run errors and Kimi K3 twelve, which reduce their flip-rate denominators to 158 of 162 disputed cases; GPT-5.6 Sol recorded one. All remain within the gate's error-rate envelope. The eight models published span the full accuracy range observed.
