AI BENCHMARK PROFILE
RedundancyBench
RedundancyBench is a benchmark for detecting redundant steps in agent trajectories. It contains diverse tasks with annotated trajectories where each step is labeled for its contribution to task completion.
- Released
- 2026-05-28
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
LLM-based agents often execute with inefficiencies, but existing evaluations focus only on task success. RedundancyBench addresses the gap in evaluating execution efficiency.
Motivation
LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.