AI BENCHMARK PROFILE
QVal
QVal is a training-free testbed for evaluating dense supervision signals for long-horizon LLM agents. It measures Q-alignment of scores from methods across four environments and seven families, without requiring training runs.
- Released
- 2026-06-30
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Fills the gap of direct, comparable evaluation of dense supervision methods, which is currently expensive and confounded. Enables early-stage assessment before training.
Motivation
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.