AAL Certified · Independent AI Audit

The independent benchmark for AI you're about to deploy.

AAL Certified tests frontier AI against real capital markets operations workflows — scored by a deterministic engine, published with full methodology, and reproducible by anyone. The same standard we apply to GPT-4o, Gemini, and Claude. Now open to external systems.

Get certified →Evaluating a vendor? See Engage →Get API key →
For AI vendors

Get certified. Close deals faster.

Enterprise capital markets clients want independent validation before signing. AAL Certified gives your sales team a third-party audit report — not a self-reported benchmark, not a vendor white paper — that stands up to CRO scrutiny.

The badge says: we didn't just test this. We published the methodology so anyone can verify the result.

For capital markets firms

Evaluating a vendor before you deploy?

Your vendor showed you a benchmark. You don't know what cases they tested, whether ground truth was pre-constructed, or what happens when the model fails.

That's a different engagement than AAL Certified — a private, bespoke assessment built around your own workflow, not a public badge. See Independent AI Assessment →

The standard

Same methodology. Every system.

Multi-run evaluation — 3 runs per case at temperature 0
Deterministic scoring — no LLM touches the scoring logic
Wilson 95% confidence intervals on every accuracy metric
Contamination audit gate — every result file passes before publication
Ground truth pre-constructed — before any model sees a case
Full disclosure — methodology, prompts, and rubrics published alongside results

This is the same standard we apply to GPT-4o, Gemini 2.5 Pro, Claude Sonnet, DeepSeek, and KIMI across 1,500 published benchmark cases. See the published results →

Sample deliverable

What an audit report looks like.

Every AAL Certified audit produces a scored, risk-committee-ready PDF. Below is the structure of a completed audit — anonymized, with all sections intact.

AAL-CERT-2026-001 · Sample Report

Independent AI Audit Report

System Under Test: TradeSync-2026 by [Vendor X] · Benchmark: AAL-D-002 · Margin Call Dispute Detection · Audit Date: 15 July 2026

DETECTION
97.2%
Threshold: ≥ 95% · Pass
ARITHMETIC
64.1%
Threshold: ≥ 60% · Pass

Executive summary. TradeSync-2026 met AAL Certified thresholds for detection and false-positive rate. Arithmetic accuracy fell within the minimum threshold but exhibited material variance on SIMM and haircut recalculation tasks.

Methodology. 250 cases across 17 dispute categories. 3 runs per case at temperature 0. Deterministic Python scoring engine. Wilson 95% confidence intervals on all proportion metrics.

Certification decision. AAL Certified — valid through 22 July 2027. Badge usage rights granted per AAL-CERT-POL-001.

Full sample report includes: System Under Test table, per-dimension breakdown, failure mode analysis with case IDs, limitation disclosures, appendices with prompt templates and scoring engine source hash. Request the full sample →

Certification tiers

For vendors. Three tiers.

Single dataset
AAL Benchmarked
$3,500one-time
One dataset (D-001, D-002, or D-003)
Full scored report
Results published on AAL research portal
No minimum threshold required
Valid 12 months
Get benchmarked
All three datasets
AAL Certified
$12,500one-time
All three datasets (D-001 + D-002 + D-003)
Must meet minimum thresholds
Detection ≥ 95% · FPR ≤ 2% · Arithmetic ≥ 60%
AAL Certified badge (digital + PDF)
Results published on AAL research portal
Valid 12 months
Get certified
Production-grade
AAL Certified: Enterprise
From $40,000
All three datasets + custom asset class extension
Elevated thresholds: detection ≥ 98%, FPR ≤ 0.5%
Risk committee PDF report
Named model version pinned and documented
Quarterly re-certification available
Private results option
Contact us
For capital markets firms
Evaluating a vendor's system before you deploy it?

That's Independent AI Assessment, not AAL Certified — a private, bespoke evaluation built around your own workflow, with a deployment recommendation for your risk committee. No badge, no publication.

How it works

Submit. Evaluate. Publish. 2 weeks.

01

Submit

Provide API credentials, model documentation, and asset class selection. We evaluate your system blind to your marketing materials.

02

Evaluate

3 runs per case at temperature 0. Deterministic scoring engine. Same methodology as our published GPT-4o, Gemini, and Claude evaluations.

03

Audit

Every result file passes our contamination audit gate. Confidence intervals computed. Methodology documented. Nothing ships without passing it.

04

Report

Scored PDF with per-dimension breakdown: detection, classification, arithmetic, escalation, false positive rate. All with Wilson 95% CIs.

05

Publish

Results published on aialphalabs.ai/research with full methodology — or kept private at your request. Turnaround: 2 weeks.

FAQ

Common questions.

What happens if my system fails the Certified thresholds?

You receive an AAL Benchmarked report with full scores and failure-mode analysis. You may retake once within 90 days at 50% of the current tier fee. If you pass on retake, the 12-month badge period begins from the retake publication date.

Are results published if I fail?

Only if you opt in. By default, failed audits remain private. However, AAL reserves the right to publish aggregate statistics (e.g., 'X% of systems tested on AAL-D-002 met the arithmetic threshold') without identifying the vendor.

Can I see my results before publication?

Yes. All vendors receive a 48-hour review window to check for factual errors in the system description. You may not alter the scores or methodology.

What happens to my API credentials?

Credentials are stored encrypted at rest (AES-256) and are deleted within 72 hours of audit completion. We do not train models on your responses. A full security memo is available under NDA.

Do you sign our vendor NDAs?

We sign mutual NDAs for Enterprise and buy-side audits. For standard AAL Certified tiers, our public terms govern the engagement.

How can we use the AAL Certified badge?

You may display the digital badge and PDF certificate on your website, in sales decks, and in RFP responses. You must use the exact language from the audit report. You may not claim certification for datasets or model versions not explicitly tested.

Can the badge be revoked?

Yes. If a vendor materially alters the certified system (e.g., model version change, prompt architecture change) without re-certification, or if AAL discovers misrepresentation of the audit scope, the badge is revoked and a public correction is issued.

What if we update our model?

The badge is tied to the pinned version. A new version requires re-certification. For teams shipping frequently, Enterprise includes quarterly re-certification at a reduced rate.

How long does the audit take?

Two weeks from credential verification to published report. Rush delivery (5 business days) is available for Enterprise clients at +30% fee.

Do you audit open-source models?

Yes. We can evaluate any system with an API endpoint that adheres to our input/output schema. Open-source weights require you to host them for our evaluation.

Can a capital markets firm audit its own internal model?

That's an Independent AI Assessment engagement (see /engage), not AAL Certified. Certified is a vendor certification product built on our published datasets with pass/fail thresholds and an optional public badge. An internal or buy-side evaluation uses the same underlying methodology but is scoped to your own workflow, stays private by default, and produces a deployment recommendation rather than a certification decision.

Who grades the audit?

AAL staff engineers run the evaluation. The contamination gate and final score verification are performed by a separate reviewer (two-person rule). The methodology is deterministic; no LLM grades the responses.

Is AAL compensated differently if we pass or fail?

No. The fee is fixed regardless of outcome. We have no equity, referral, or success-based arrangements with any model vendor.

Independence

We have no commercial interest in your result.

We don't build the models we test. We don't sell the AI we evaluate. We don't take commissions from model vendors.

The only product is the evidence. Think of it like a rating agency — the value of the rating is the independence.

1,500 benchmark cases, six datasets, full methodology — all published and reproducible. See the research →

Ready to get started? Email us to discuss certification or commission an audit. We respond same day.

joe@aialphalabs.ai