Independent AI assessment for capital markets operations.
“Your vendor says it's 99% accurate.” On which cases? How many runs? Scored by whom?
We independently evaluate AI systems on the operations workflows your desk actually runs — trade confirmation, margin disputes, settlement fails, P&L reconciliation, collateral eligibility. We measure the system against ground truth established before evaluation begins, score it with a deterministic engine, and publish the methodology so anyone can reproduce the result.
We don't build the models we test. We don't sell the AI we evaluate. The only product is the evidence.
Across six published benchmarks. Models reliably know something is wrong; they cannot reliably tell you by how much. A single “the model is validated” sign-off hides exactly that risk.
SR 26-2 (Fed / OCC / FDIC, April 2026) replaced SR 11-7 and left standalone generative and agentic AI outside the formal model-risk framework — putting validation design back on the institution, with no supervisory template for what sufficient evidence looks like.
The SEC's 2026 examination prioritiesname AI governance directly. Examiners are reviewing the accuracy of firms' AI-capability representations, not merely whether a policy exists.
“The vendor told us” is no longer an answer. Independent, reproducible evidence is.
Two questions most evaluations conflate.
Detection, classification, attribution, false-positive rate.
Value accuracy, exposure calculation, escalation judgment.
These fail at completely different rates. Measuring them together hides the failure that matters.
A defensible evidence package, not a slide deck.
Every metric with Wilson 95% confidence intervals.
Full scored output — every case, every run.
Prompts, rubrics, and scoring logic, versioned and open.
What a third party needs to independently re-run the evaluation.
Which tasks the system is fit for, which require deterministic controls, and where a human must sign.
Written for model risk, audit committee, or examiner review.
Worked example — this is what your deliverable looks like. AAL-D-006, Collateral Eligibility & Substitution: 250 cases, nine frontier models, 6,750 scored observations, published in full. View AAL-D-006 →
Six to eight weeks, end to end.
Scoping
Week 1
Workflow selection, success criteria, and metrics agreed up front.
Dataset construction
Weeks 2–3
Cases built with ground truth by construction, then independently recomputed before anything is scored.
Evaluation
Weeks 4–5
Multi-run execution, deterministic scoring, and a pre-publication integrity gate.
Reporting
Weeks 6–7
Scorecards, findings, deployment recommendation, and a review session.
No access to production systems required. Evaluations run on constructed cases that mirror your workflows, so nothing sensitive leaves your environment.
Fixed scope. Fixed fee. Paid regardless of outcome.
The fee does not change based on what we find. We are paid for the measurement, not the verdict — and unfavorable findings are reported as clearly as favorable ones. That is the entire point.
Selling AI and need a publishable, pass/fail badge instead of a private assessment? See AAL Certified →
$35,000
- ·One operations workflow
- ·Up to three models evaluated
- ·Three runs per case, deterministic scoring
- ·Full evidence package
- ·6–8 weeks
$55,000
- ·One operations workflow
- ·Up to eight models evaluated
- ·Head-to-head comparison across the field
- ·Full evidence package
- ·6–8 weeks
$60,000
- ·New dataset built to your workflow
- ·Generalized version published to the public series
- ·You fund the measurement; the field gets the standard
- ·Full evidence package
- ·8–10 weeks
No model vendor relationships, no AI products sold, and no consulting to help anyone pass a test we administer.
Six datasets, 1,500 validated cases, ten frontier models, deterministic scoring — everything published, and the reference set your results are calibrated against.
Fifteen years in capital markets operations across Millennium, Schonfeld, Fannie Mae and Vanguard. The workflows are modeled by someone who ran them.
Start with the evidence.
Before you buy from any AI vendor, send them the 10-question Reproducibility Scorecard. If they score below 5, their claims aren't independently verifiable — and that is worth knowing before a procurement decision, not after.
