Benchmark Radar
AI BENCHMARK PROFILE

RedundancyBench

General AIKnowledge & Reasoning

RedundancyBench is a benchmark for detecting redundant steps in agent trajectories. It contains diverse tasks with annotated trajectories where each step is labeled for its contribution to task completion.

Released
2026-05-28
Readiness
Inspectable
Primary field
General AI

Why it matters

LLM-based agents often execute with inefficiencies, but existing evaluations focus only on task success. RedundancyBench addresses the gap in evaluating execution efficiency.

Motivation

LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.