Engage AI Alpha Labs

Independent AI assessment for capital markets operations.

“Your vendor says it's 99% accurate.” On which cases? How many runs? Scored by whom?

We independently evaluate AI systems on the operations workflows your desk actually runs — trade confirmation, margin disputes, settlement fails, P&L reconciliation, collateral eligibility. We measure the system against ground truth established before evaluation begins, score it with a deterministic engine, and publish the methodology so anyone can reproduce the result.

We don't build the models we test. We don't sell the AI we evaluate. The only product is the evidence.

93–100%
Detection accuracy
12–58%
Value accuracy
1,500
Validated cases
10
Frontier models

Across six published benchmarks. Models reliably know something is wrong; they cannot reliably tell you by how much. A single “the model is validated” sign-off hides exactly that risk.

Why now

SR 26-2 (Fed / OCC / FDIC, April 2026) replaced SR 11-7 and left standalone generative and agentic AI outside the formal model-risk framework — putting validation design back on the institution, with no supervisory template for what sufficient evidence looks like.

The SEC's 2026 examination prioritiesname AI governance directly. Examiners are reviewing the accuracy of firms' AI-capability representations, not merely whether a policy exists.

“The vendor told us” is no longer an answer. Independent, reproducible evidence is.

What we measure

Two questions most evaluations conflate.

01
Does the system know a problem exists?

Detection, classification, attribution, false-positive rate.

93–100%
02
Can it compute the right answer?

Value accuracy, exposure calculation, escalation judgment.

12–58%

These fail at completely different rates. Measuring them together hides the failure that matters.

What you receive

A defensible evidence package, not a slide deck.

01
Per-model scorecard

Every metric with Wilson 95% confidence intervals.

02
Case-level results

Full scored output — every case, every run.

03
Published methodology

Prompts, rubrics, and scoring logic, versioned and open.

04
Reproducibility statement

What a third party needs to independently re-run the evaluation.

05
Deployment recommendation

Which tasks the system is fit for, which require deterministic controls, and where a human must sign.

06
Regulator-ready summary

Written for model risk, audit committee, or examiner review.

Worked example — this is what your deliverable looks like. AAL-D-006, Collateral Eligibility & Substitution: 250 cases, nine frontier models, 6,750 scored observations, published in full. View AAL-D-006 →

How it works

Six to eight weeks, end to end.

01

Scoping

Week 1

Workflow selection, success criteria, and metrics agreed up front.

02

Dataset construction

Weeks 2–3

Cases built with ground truth by construction, then independently recomputed before anything is scored.

03

Evaluation

Weeks 4–5

Multi-run execution, deterministic scoring, and a pre-publication integrity gate.

04

Reporting

Weeks 6–7

Scorecards, findings, deployment recommendation, and a review session.

No access to production systems required. Evaluations run on constructed cases that mirror your workflows, so nothing sensitive leaves your environment.

Engagement terms

Fixed scope. Fixed fee. Paid regardless of outcome.

The fee does not change based on what we find. We are paid for the measurement, not the verdict — and unfavorable findings are reported as clearly as favorable ones. That is the entire point.

Selling AI and need a publishable, pass/fail badge instead of a private assessment? See AAL Certified →

Single workflowUp to 3 models

$35,000

  • ·One operations workflow
  • ·Up to three models evaluated
  • ·Three runs per case, deterministic scoring
  • ·Full evidence package
  • ·6–8 weeks
Multi-model comparisonUp to 8 models

$55,000

  • ·One operations workflow
  • ·Up to eight models evaluated
  • ·Head-to-head comparison across the field
  • ·Full evidence package
  • ·6–8 weeks
Commissioned benchmarkNew dataset

$60,000

  • ·New dataset built to your workflow
  • ·Generalized version published to the public series
  • ·You fund the measurement; the field gets the standard
  • ·Full evidence package
  • ·8–10 weeks
Why AI Alpha Labs
01
Independent by construction

No model vendor relationships, no AI products sold, and no consulting to help anyone pass a test we administer.

02
The only public corpus of its kind

Six datasets, 1,500 validated cases, ten frontier models, deterministic scoring — everything published, and the reference set your results are calibrated against.

03
Operator-built

Fifteen years in capital markets operations across Millennium, Schonfeld, Fannie Mae and Vanguard. The workflows are modeled by someone who ran them.

Start with the evidence.

Before you buy from any AI vendor, send them the 10-question Reproducibility Scorecard. If they score below 5, their claims aren't independently verifiable — and that is worth knowing before a procurement decision, not after.