Benchmark Radar
AI BENCHMARK PROFILE

AgentLens

General AIAgents

AgentLens evaluates interactive coding agents across entire trajectories, combining formal verification with LLM-written trajectory reviews and side-by-side comparisons to score dimensions such as instruction compliance, tool use, and interaction style.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Most code-agent benchmarks reduce an episode to a single pass/fail bit, which is too coarse for production use. AgentLens provides a richer evaluation that captures real user experience and yields readable justifications for scores, aiding regression detection and model diagnosis.

Motivation

We present AgentLens, a production-assessed benchmark for interactive code agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.