← Research Portal
AAL-M-2026-001 v1.0

The AAL Evaluation Protocol

Seven principles. Nine phases. Deterministic scoring. This is the methodology that governs every research publication, benchmark evaluation, and lab validation conducted by AI Alpha Labs.

Seven principles

These are not aspirational — they are requirements. An evaluation that violates any principle is non-conformant. Each principle links to the published work that demonstrates it.

4.1
Independence

Evaluations are conducted independently of model providers. No provider reviews, approves, or influences methodology, scoring, or findings.

AAL-R-2026-001
4.2
Transparency

All methodology, scoring criteria, confidence levels, limitations, and failure modes are disclosed. No element of the evaluation is withheld from the reader.

AAL-D-001
4.3
Reproducibility

Sufficient information is published for an independent party to replicate the evaluation. A finding that cannot be reproduced is not a finding.

AAL-D-002
4.4
Evidence basis

Findings are derived from data. Conclusions do not exceed what the evidence supports. Qualitative assessments are labeled as such.

AAL-F-001
4.5
Operational grounding

Evaluations are conducted on tasks derived from real institutional workflows. General-purpose or academic tasks are not used.

AAL-D-003
4.6
Bias mitigation

All scoring is conducted blind. The identity of the model producing each output is concealed from reviewers until scoring is complete.

AAL-M-2026-001
4.7
Limitation disclosure

Known limitations are documented before findings are drafted. Limitations are published with equal prominence to findings.

AAL-F-002

Nine-phase lifecycle

Every evaluation follows this sequence. Each phase produces defined artifacts. No phase may be omitted.

1

Research Design

Research question · objectives · scope · workflow selection

2

Data Collection

Dataset sourcing · anonymization · validation

3

Ground Truth Construction

Expert authoring · independent review · consensus

4

Human Analyst Baseline

Qualified analysts · identical tasks · independent scoring

5

Model Evaluation

Prompt lock · model execution · output capture

6

Blind Review and Scoring

Blind assembly · independent scoring · adjudication

7

Statistical Analysis

Metrics · confidence intervals · classifications

8

Limitations and Threats to Validity

Internal · external · construct · documented before findings

9

Publication

Findings · internal review · version · release

Terms and definitions

From AAL-M-2026-001 §3. Defined terms are used consistently across all publications.

Accuracy

The proportion of tasks where the model's output matches ground truth on all required fields, as determined by qualified blind reviewers.

Blind review

An evaluation procedure in which the reviewer does not know which model produced the output being scored.

Confidence score

A quantitative measure of the statistical robustness of an evaluation's findings. Confidence reflects the quality of the evidence, not the performance of any model.

Fabrication

Any output element not present in, or directly derivable from, the input provided to the model. Synonymous with 'hallucination' in colloquial usage.

Ground truth

The correct output for a given task, established by qualified domain experts before model evaluation begins.

Human analyst baseline

The measured performance of qualified operations professionals completing identical tasks without AI assistance, scored using the same rubric applied to model outputs.

Production-capable

A model that meets all evaluation thresholds simultaneously for a given workflow.

Scoring rubric

A predefined set of criteria that determines how model outputs are scored against ground truth. Rubrics are fixed before evaluation begins.

Workflow

A category of operational work evaluated under this methodology (e.g., trade confirmation matching, reconciliation break analysis).

Normative documents

AAL-M-2026-001
Research Methodology

Governs every research publication, benchmark evaluation, and lab validation. Defines experimental design, data collection, ground truth construction, model evaluation, blind review, statistical analysis, and publication.

v1.0
AAL-S-2026-001
Standards

Defines requirements for datasets, evaluations, research publications, benchmarks, results, confidence calculations, version control, citation practices, and review processes.

v1.0
AAL-D-2026-001
Dataset Framework

Defines task structure, ground truth requirements, anonymization rules, and validation criteria for all AAL datasets.

Framework

Want to audit our methodology? All normative documents are published on GitHub. Every dataset, prompt, and scorer is available for independent replication.

View documents on GitHub →