The independent benchmark for AI you're about to deploy.
AAL Certified tests frontier AI against real capital markets operations workflows — scored by a deterministic engine, published with full methodology, and reproducible by anyone. The same standard we apply to GPT-4o, Gemini, and Claude. Now open to external systems.
Get certified. Close deals faster.
Enterprise capital markets clients want independent validation before signing. AAL Certified gives your sales team a third-party audit report — not a self-reported benchmark, not a vendor white paper — that stands up to CRO scrutiny.
The badge says: we didn't just test this. We published the methodology so anyone can verify the result.
Evaluating a vendor before you deploy?
Your vendor showed you a benchmark. You don't know what cases they tested, whether ground truth was pre-constructed, or what happens when the model fails.
That's a different engagement than AAL Certified — a private, bespoke assessment built around your own workflow, not a public badge. See Independent AI Assessment →
Same methodology. Every system.
This is the same standard we apply to GPT-4o, Gemini 2.5 Pro, Claude Sonnet, DeepSeek, and KIMI across 1,500 published benchmark cases. See the published results →
What an audit report looks like.
Every AAL Certified audit produces a scored, risk-committee-ready PDF. Below is the structure of a completed audit — anonymized, with all sections intact.
Independent AI Audit Report
System Under Test: TradeSync-2026 by [Vendor X] · Benchmark: AAL-D-002 · Margin Call Dispute Detection · Audit Date: 15 July 2026
Executive summary. TradeSync-2026 met AAL Certified thresholds for detection and false-positive rate. Arithmetic accuracy fell within the minimum threshold but exhibited material variance on SIMM and haircut recalculation tasks.
Methodology. 250 cases across 17 dispute categories. 3 runs per case at temperature 0. Deterministic Python scoring engine. Wilson 95% confidence intervals on all proportion metrics.
Certification decision. AAL Certified — valid through 22 July 2027. Badge usage rights granted per AAL-CERT-POL-001.
Full sample report includes: System Under Test table, per-dimension breakdown, failure mode analysis with case IDs, limitation disclosures, appendices with prompt templates and scoring engine source hash. Request the full sample →
For vendors. Three tiers.
That's Independent AI Assessment, not AAL Certified — a private, bespoke evaluation built around your own workflow, with a deployment recommendation for your risk committee. No badge, no publication.
Submit. Evaluate. Publish. 2 weeks.
Submit
Provide API credentials, model documentation, and asset class selection. We evaluate your system blind to your marketing materials.
Evaluate
3 runs per case at temperature 0. Deterministic scoring engine. Same methodology as our published GPT-4o, Gemini, and Claude evaluations.
Audit
Every result file passes our contamination audit gate. Confidence intervals computed. Methodology documented. Nothing ships without passing it.
Report
Scored PDF with per-dimension breakdown: detection, classification, arithmetic, escalation, false positive rate. All with Wilson 95% CIs.
Publish
Results published on aialphalabs.ai/research with full methodology — or kept private at your request. Turnaround: 2 weeks.
Common questions.
What happens if my system fails the Certified thresholds?
You receive an AAL Benchmarked report with full scores and failure-mode analysis. You may retake once within 90 days at 50% of the current tier fee. If you pass on retake, the 12-month badge period begins from the retake publication date.
Are results published if I fail?
Only if you opt in. By default, failed audits remain private. However, AAL reserves the right to publish aggregate statistics (e.g., 'X% of systems tested on AAL-D-002 met the arithmetic threshold') without identifying the vendor.
Can I see my results before publication?
Yes. All vendors receive a 48-hour review window to check for factual errors in the system description. You may not alter the scores or methodology.
What happens to my API credentials?
Credentials are stored encrypted at rest (AES-256) and are deleted within 72 hours of audit completion. We do not train models on your responses. A full security memo is available under NDA.
Do you sign our vendor NDAs?
We sign mutual NDAs for Enterprise and buy-side audits. For standard AAL Certified tiers, our public terms govern the engagement.
How can we use the AAL Certified badge?
You may display the digital badge and PDF certificate on your website, in sales decks, and in RFP responses. You must use the exact language from the audit report. You may not claim certification for datasets or model versions not explicitly tested.
Can the badge be revoked?
Yes. If a vendor materially alters the certified system (e.g., model version change, prompt architecture change) without re-certification, or if AAL discovers misrepresentation of the audit scope, the badge is revoked and a public correction is issued.
What if we update our model?
The badge is tied to the pinned version. A new version requires re-certification. For teams shipping frequently, Enterprise includes quarterly re-certification at a reduced rate.
How long does the audit take?
Two weeks from credential verification to published report. Rush delivery (5 business days) is available for Enterprise clients at +30% fee.
Do you audit open-source models?
Yes. We can evaluate any system with an API endpoint that adheres to our input/output schema. Open-source weights require you to host them for our evaluation.
Can a capital markets firm audit its own internal model?
That's an Independent AI Assessment engagement (see /engage), not AAL Certified. Certified is a vendor certification product built on our published datasets with pass/fail thresholds and an optional public badge. An internal or buy-side evaluation uses the same underlying methodology but is scoped to your own workflow, stays private by default, and produces a deployment recommendation rather than a certification decision.
Who grades the audit?
AAL staff engineers run the evaluation. The contamination gate and final score verification are performed by a separate reviewer (two-person rule). The methodology is deterministic; no LLM grades the responses.
Is AAL compensated differently if we pass or fail?
No. The fee is fixed regardless of outcome. We have no equity, referral, or success-based arrangements with any model vendor.
We have no commercial interest in your result.
We don't build the models we test. We don't sell the AI we evaluate. We don't take commissions from model vendors.
The only product is the evidence. Think of it like a rating agency — the value of the rating is the independence.
1,500 benchmark cases, six datasets, full methodology — all published and reproducible. See the research →
Ready to get started? Email us to discuss certification or commission an audit. We respond same day.
joe@aialphalabs.ai