AI BENCHMARK PROFILE
GAUGE
GAUGE evaluates agent-built financial valuation models against observed analyst practice using 56 facets, eight validity gates, and a failure-aware score over 196 tasks.
- Released
- 2026-07-27
- Readiness
- Paper only
- Primary field
- Finance & Economics
Why it matters
Provides a benchmark that avoids penalizing legitimate disagreement, enabling fair comparison of agents on financial modeling and judgment.
Motivation
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.