AI BENCHMARK PROFILE
AgentLens
AgentLens evaluates interactive coding agents across entire trajectories, combining formal verification with LLM-written trajectory reviews and side-by-side comparisons to score dimensions such as instruction compliance, tool use, and interaction style.
- Released
- 2026-07-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Most code-agent benchmarks reduce an episode to a single pass/fail bit, which is too coarse for production use. AgentLens provides a richer evaluation that captures real user experience and yields readable justifications for scores, aiding regression detection and model diagnosis.
Motivation
We present AgentLens, a production-assessed benchmark for interactive code agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.