← Research Portal
AAL-F-008Confirmed

A 30B open-weight model matches frontier detection — and hits the same value wall

Findings·Aug 2026
95.7%
Detection
78.6%
Substitution validity
16.3%
Value accuracy
Single consumer GPU, local, open-weight
Deployment
100% / 92.6% / 31.1%
Comparable closed model (Spark 1.2)

Run against the AAL-D-006 collateral eligibility and substitution set (250 cases, 3 runs, 0 errors), Muse Glimmer — Meta's 30B open-weight model, deployable locally on a single consumer GPU — outperforms DeepSeek V4 Pro on detection (95.7% vs 93.1%) and outperforms both Sonnet 5 (68.1%) and DeepSeek V4 Pro (66.9%) on substitution validity (78.6%). Its value accuracy of 16.3% sits within the same 15–58% band every frontier model in the corpus occupies. Held against Spark 1.2 — the closed, frontier-tier version of the same model family — the pattern holds inside a single vendor's own lineup: Spark 1.2 scores higher across the board (100% detection, 92.6% substitution, 31.1% value) but hits an equivalent wall on value. Not published in the flagship comparison table: Glimmer's 30B parameter count makes a head-to-head against 200B+ closed models a size-mismatched comparison. This entry isolates the finding it actually supports.

The wall is not a scale problem

A 30-billion-parameter model that runs inside a bank's own walls matches the frontier field on detection and beats two flagships on substitution — then lands on the identical quantification wall as models 10x its size. Scale buys detection and judgment. It does not buy the number.

The on-prem case

For regulated firms that can't send collateral schedules to a vendor API, Glimmer-class local models are a realistic path for the triage layer: detect, categorize, propose a substitute, all on hardware already inside the building, with an open-weight model anyone can independently reproduce results against. Valuation still has to route to deterministic code — no model in the corpus, open or closed, computes the number reliably enough to book.

Methodology

Same AAL-D-006 dataset, prompts, and deterministic scorer as the flagship comparison — 250 cases, 3 runs, 0 errors, temp 0.7 / top-p 0.9. Not blended into the flagship table: comparing a 30B open-weight model against 200B+ closed frontier models on a single ranked table would overstate the significance of any gap in either direction.

Back to portalView full series →
Related