Benchmark Radar
AI BENCHMARK PROFILE

QVal

General AIKnowledge & Reasoningbethgelab

QVal is a training-free testbed for evaluating dense supervision signals for long-horizon LLM agents. It measures Q-alignment of scores from methods across four environments and seven families, without requiring training runs.

Released
2026-06-30
Readiness
Runnable
Primary field
General AI

Why it matters

Fills the gap of direct, comparable evaluation of dense supervision methods, which is currently expensive and confounded. Enables early-stage assessment before training.

Motivation

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.