Benchmark Radar
AI BENCHMARK PROFILE

Agent Retrieval Bench

General AIKnowledge & ReasoningSearch & RetrievalCoding & Software EngineeringAgent Retrieval Bench team

File-level retrieval benchmark for coding agents, covering four positive tasks (code2test, comment2context, trace2code, edit2ripple) and a selective-retrieval subset with natural no-gold and counterfactual controls across 25 repositories, 427 samples, with frozen base-commit corpora.

Released
2026-07-27
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a dedicated evaluation for the context-acquisition stage of coding agents, distinguishing retrieval quality from patch generation and offering a reusable protocol for comparing retrieval and selective abstention methods.

Motivation

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.