A pre-qualification questionnaire for evaluating AI vendor benchmark claims in capital markets operations.
Most AI accuracy figures cited in capital markets procurement are neither reproducible nor independently validated. Send this scorecard to any AI vendor before accuracy claims are discussed. A vendor declining to complete it should be disqualified from further consideration.
| # | Question | Vendor Response | Score |
|---|---|---|---|
| 1 | Model version disclosureState the exact model version(s) used (e.g., gpt-4o-2024-08-06, claude-sonnet-5). List all if multiple. | — | ☐ /1 |
| 2 | Dataset size & compositionTotal cases evaluated, and stratification by asset class, workflow, and difficulty. | — | ☐ /1 |
| 3 | Independent runsRuns per case and the sampling protocol (temperature and any variation). | — | ☐ /1 |
| 4 | Cross-model validationWere multiple frontier models run against the same cases? List them; describe disagreement handling. | — | ☐ /1 |
| 5 | Published evaluation rubricAttach or link the full rubric: field definitions, scoring rules, edge-case handling. | — | ☐ /1 |
| 6 | Data provenanceProduction, synthetic, or vendor-created? If production, confirm separation from training data. | — | ☐ /1 |
| 7 | Leakage controlConfirm no test case appeared in any training or tuning set; describe the separation or construction protocol. | — | ☐ /1 |
| 8 | Arithmetic-validation architectureHow are numerical claims validated? If the LLM does arithmetic, how is it verified? If code does, describe the handoff. | — | ☐ /1 |
| 9 | Marginal cost disclosureApproximate API cost to reproduce one evaluation pass at current pricing, with model(s) and tokens per case. | — | ☐ /1 |
| 10 | Independent replicationHas a third party replicated the benchmark? If yes: name, date, findings. If no: state your external-replication policy. | — | ☐ /1 |
| TOTAL | /10 |
We hold ourselves to the standard we ask of vendors. Scored against our own benchmark corpus (AAL-D-001 through D-006).
| # | Criterion | AAL Response | Score |
|---|---|---|---|
| 1 | Model version disclosure | Roster disclosed per dataset. Version strings include gpt-4o-2024-08-06, gemini-2.5-pro, gemini-2.5-flash, claude-sonnet-4-6, claude-sonnet-5, gpt-5.6-sol, claude-opus-5, gemini-3.1-pro-preview, x-ai/grok-4.5, deepseek/deepseek-v4-pro, moonshotai/kimi-k3, meta/muse-spark-1.2, qwen3.8-max. D-006 uses a fixed nine-model frontier roster (plus a standalone open-weight addendum, Muse Glimmer 30B, reported separately). | 1 |
| 2 | Dataset size & composition | 1,500 cases across six datasets. D-001: 250 (7 asset classes). D-002: 250 (17 categories). D-003: 250 (19 categories). D-004: 250 (162 fails / 88 trap cleans). D-005: 250 (162 breaks / 88 explained-move traps). D-006: 250 (162 ineligible / 88 traps, 12 eligibility categories). Stratified by difficulty. | 1 |
| 3 | Independent runs | Three runs per case, temperature 0. Reasoning-model families run at deterministic default. Pooled Wilson 95% CIs on all figures. | 1 |
| 4 | Cross-model validation | Up to nine frontier models on identical cases (6,000 obs on D-003; 4,500 on D-005; 6,750 on D-006). Cross-vendor comparison reported across ten models to date; disagreements surfaced. | 1 |
| 5 | Published evaluation rubric | Deterministic, criteria-keyed scorer with per-case scoring_criteria; full dataset specifications published. | 1 |
| 6 | Data provenance | Synthetic, ground-truth-by-construction from locked specifications — not client production data. Fully disclosable; datasets, prompts, scorers published post-evaluation. | 1 |
| 7 | Leakage control | No leakage path exists — cases derive from no real corpus. Pre-publication SHA-256 registration; no-leakage grep confirms no ground-truth string in any serialized prompt. | 1 |
| 8 | Arithmetic-validation architecture | Deterministic Python performs all arithmetic; LLMs classify only. No model performs a day-count, exposure, call amount, or residual. Non-negotiable across all datasets. | 1 |
| 9 | Marginal cost disclosure | ~$300–400 for the full corpus (1,500 cases × up to nine models × three runs) — cents per case. Per-dataset costs disclosed in AAL-RS-009. | 1 |
| 10 | Independent replication | Not yet independently replicated. Specifications, prompts, and scripts published to enable it; treated as the outstanding step. | 0.5 |
| AAL TOTAL | 0.5 deducted for pending independent replication. | 9.5 / 10 |
Complete and return with any vendor proposal. Incomplete scorecards should be returned for revision; vendors scoring below 5 should not advance to subsequent evaluation rounds. Verify claims by (1) requesting the prompt templates, (2) running a 50-case subset through the vendor's stated protocol and disclosed model versions, and (3) comparing the vendor's reported accuracy against your independent run. A claim that cannot survive this verification should not survive contract negotiation.