AI Alpha Labs Independent AI Evaluation · Capital Markets
Reproducibility Scorecard · v1.0 · AAL-RS-010

The AAL Reproducibility Scorecard

A pre-qualification questionnaire for evaluating AI vendor benchmark claims in capital markets operations.

Most AI accuracy figures cited in capital markets procurement are neither reproducible nor independently validated. Send this scorecard to any AI vendor before accuracy claims are discussed. A vendor declining to complete it should be disqualified from further consideration.

Scoring. Each question scores Yes = 1, Partial = 0.5, or No = 0. Total possible: 10 points.
9–10Exceptional transparency. The vendor treats benchmarking as evidence, not marketing. 7–8Adequate. Genuine effort with gaps to close before contract execution. 5–6Marginal. Minimum standards met; reliability may be overstated. Below 5Inadequate. Claims are not independently verifiable. Proceed only with significant risk mitigation.
#QuestionVendor ResponseScore
1Model version disclosureState the exact model version(s) used (e.g., gpt-4o-2024-08-06, claude-sonnet-5). List all if multiple.☐ /1
2Dataset size & compositionTotal cases evaluated, and stratification by asset class, workflow, and difficulty.☐ /1
3Independent runsRuns per case and the sampling protocol (temperature and any variation).☐ /1
4Cross-model validationWere multiple frontier models run against the same cases? List them; describe disagreement handling.☐ /1
5Published evaluation rubricAttach or link the full rubric: field definitions, scoring rules, edge-case handling.☐ /1
6Data provenanceProduction, synthetic, or vendor-created? If production, confirm separation from training data.☐ /1
7Leakage controlConfirm no test case appeared in any training or tuning set; describe the separation or construction protocol.☐ /1
8Arithmetic-validation architectureHow are numerical claims validated? If the LLM does arithmetic, how is it verified? If code does, describe the handoff.☐ /1
9Marginal cost disclosureApproximate API cost to reproduce one evaluation pass at current pricing, with model(s) and tokens per case.☐ /1
10Independent replicationHas a third party replicated the benchmark? If yes: name, date, findings. If no: state your external-replication policy.☐ /1
TOTAL/10
AI Alpha Labs Independent AI Evaluation · Capital Markets
Reproducibility Scorecard · Demonstration

AAL Self-Assessment

We hold ourselves to the standard we ask of vendors. Scored against our own benchmark corpus (AAL-D-001 through D-006).

#CriterionAAL ResponseScore
1Model version disclosureRoster disclosed per dataset. Version strings include gpt-4o-2024-08-06, gemini-2.5-pro, gemini-2.5-flash, claude-sonnet-4-6, claude-sonnet-5, gpt-5.6-sol, claude-opus-5, gemini-3.1-pro-preview, x-ai/grok-4.5, deepseek/deepseek-v4-pro, moonshotai/kimi-k3, meta/muse-spark-1.2, qwen3.8-max. D-006 uses a fixed nine-model frontier roster (plus a standalone open-weight addendum, Muse Glimmer 30B, reported separately).1
2Dataset size & composition1,500 cases across six datasets. D-001: 250 (7 asset classes). D-002: 250 (17 categories). D-003: 250 (19 categories). D-004: 250 (162 fails / 88 trap cleans). D-005: 250 (162 breaks / 88 explained-move traps). D-006: 250 (162 ineligible / 88 traps, 12 eligibility categories). Stratified by difficulty.1
3Independent runsThree runs per case, temperature 0. Reasoning-model families run at deterministic default. Pooled Wilson 95% CIs on all figures.1
4Cross-model validationUp to nine frontier models on identical cases (6,000 obs on D-003; 4,500 on D-005; 6,750 on D-006). Cross-vendor comparison reported across ten models to date; disagreements surfaced.1
5Published evaluation rubricDeterministic, criteria-keyed scorer with per-case scoring_criteria; full dataset specifications published.1
6Data provenanceSynthetic, ground-truth-by-construction from locked specifications — not client production data. Fully disclosable; datasets, prompts, scorers published post-evaluation.1
7Leakage controlNo leakage path exists — cases derive from no real corpus. Pre-publication SHA-256 registration; no-leakage grep confirms no ground-truth string in any serialized prompt.1
8Arithmetic-validation architectureDeterministic Python performs all arithmetic; LLMs classify only. No model performs a day-count, exposure, call amount, or residual. Non-negotiable across all datasets.1
9Marginal cost disclosure~$300–400 for the full corpus (1,500 cases × up to nine models × three runs) — cents per case. Per-dataset costs disclosed in AAL-RS-009.1
10Independent replicationNot yet independently replicated. Specifications, prompts, and scripts published to enable it; treated as the outstanding step.0.5
AAL TOTAL0.5 deducted for pending independent replication.9.5 / 10

How to use this scorecard

Complete and return with any vendor proposal. Incomplete scorecards should be returned for revision; vendors scoring below 5 should not advance to subsequent evaluation rounds. Verify claims by (1) requesting the prompt templates, (2) running a 50-case subset through the vendor's stated protocol and disclosed model versions, and (3) comparing the vendor's reported accuracy against your independent run. A claim that cannot survive this verification should not survive contract negotiation.