A 30B open-weight model matches frontier detection — and hits the same value wall
Run against the AAL-D-006 collateral eligibility and substitution set (250 cases, 3 runs, 0 errors), Muse Glimmer — Meta's 30B open-weight model, deployable locally on a single consumer GPU — outperforms DeepSeek V4 Pro on detection (95.7% vs 93.1%) and outperforms both Sonnet 5 (68.1%) and DeepSeek V4 Pro (66.9%) on substitution validity (78.6%). Its value accuracy of 16.3% sits within the same 15–58% band every frontier model in the corpus occupies. Held against Spark 1.2 — the closed, frontier-tier version of the same model family — the pattern holds inside a single vendor's own lineup: Spark 1.2 scores higher across the board (100% detection, 92.6% substitution, 31.1% value) but hits an equivalent wall on value. Not published in the flagship comparison table: Glimmer's 30B parameter count makes a head-to-head against 200B+ closed models a size-mismatched comparison. This entry isolates the finding it actually supports.
The wall is not a scale problem
A 30-billion-parameter model that runs inside a bank's own walls matches the frontier field on detection and beats two flagships on substitution — then lands on the identical quantification wall as models 10x its size. Scale buys detection and judgment. It does not buy the number.
The on-prem case
For regulated firms that can't send collateral schedules to a vendor API, Glimmer-class local models are a realistic path for the triage layer: detect, categorize, propose a substitute, all on hardware already inside the building, with an open-weight model anyone can independently reproduce results against. Valuation still has to route to deterministic code — no model in the corpus, open or closed, computes the number reliably enough to book.
Methodology
Same AAL-D-006 dataset, prompts, and deterministic scorer as the flagship comparison — 250 cases, 3 runs, 0 errors, temp 0.7 / top-p 0.9. Not blended into the flagship table: comparing a 30B open-weight model against 200B+ closed frontier models on a single ranked table would overstate the significance of any gap in either direction.
