The AAL Evaluation Protocol
Seven principles. Nine phases. Deterministic scoring. This is the methodology that governs every research publication, benchmark evaluation, and lab validation conducted by AI Alpha Labs.
Seven principles
These are not aspirational — they are requirements. An evaluation that violates any principle is non-conformant. Each principle links to the published work that demonstrates it.
Evaluations are conducted independently of model providers. No provider reviews, approves, or influences methodology, scoring, or findings.
All methodology, scoring criteria, confidence levels, limitations, and failure modes are disclosed. No element of the evaluation is withheld from the reader.
Sufficient information is published for an independent party to replicate the evaluation. A finding that cannot be reproduced is not a finding.
Findings are derived from data. Conclusions do not exceed what the evidence supports. Qualitative assessments are labeled as such.
Evaluations are conducted on tasks derived from real institutional workflows. General-purpose or academic tasks are not used.
All scoring is conducted blind. The identity of the model producing each output is concealed from reviewers until scoring is complete.
Known limitations are documented before findings are drafted. Limitations are published with equal prominence to findings.
Nine-phase lifecycle
Every evaluation follows this sequence. Each phase produces defined artifacts. No phase may be omitted.
Research Design
Research question · objectives · scope · workflow selection
Data Collection
Dataset sourcing · anonymization · validation
Ground Truth Construction
Expert authoring · independent review · consensus
Human Analyst Baseline
Qualified analysts · identical tasks · independent scoring
Model Evaluation
Prompt lock · model execution · output capture
Blind Review and Scoring
Blind assembly · independent scoring · adjudication
Statistical Analysis
Metrics · confidence intervals · classifications
Limitations and Threats to Validity
Internal · external · construct · documented before findings
Publication
Findings · internal review · version · release
Terms and definitions
From AAL-M-2026-001 §3. Defined terms are used consistently across all publications.
The proportion of tasks where the model's output matches ground truth on all required fields, as determined by qualified blind reviewers.
An evaluation procedure in which the reviewer does not know which model produced the output being scored.
A quantitative measure of the statistical robustness of an evaluation's findings. Confidence reflects the quality of the evidence, not the performance of any model.
Any output element not present in, or directly derivable from, the input provided to the model. Synonymous with 'hallucination' in colloquial usage.
The correct output for a given task, established by qualified domain experts before model evaluation begins.
The measured performance of qualified operations professionals completing identical tasks without AI assistance, scored using the same rubric applied to model outputs.
A model that meets all evaluation thresholds simultaneously for a given workflow.
A predefined set of criteria that determines how model outputs are scored against ground truth. Rubrics are fixed before evaluation begins.
A category of operational work evaluated under this methodology (e.g., trade confirmation matching, reconciliation break analysis).
Normative documents
Governs every research publication, benchmark evaluation, and lab validation. Defines experimental design, data collection, ground truth construction, model evaluation, blind review, statistical analysis, and publication.
Defines requirements for datasets, evaluations, research publications, benchmarks, results, confidence calculations, version control, citation practices, and review processes.
Defines task structure, ground truth requirements, anonymization rules, and validation criteria for all AAL datasets.
Want to audit our methodology? All normative documents are published on GitHub. Every dataset, prompt, and scorer is available for independent replication.
View documents on GitHub →