Agent Retrieval Bench
File-level retrieval benchmark for coding agents, covering four positive tasks (code2test, comment2context, trace2code, edit2ripple) and a selective-retrieval subset with natural no-gold and counterfactual controls across 25 repositories, 427 samples, with frozen base-commit corpora.
- Released
- 2026-07-27
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a dedicated evaluation for the context-acquisition stage of coding agents, distinguishing retrieval quality from patch generation and offering a reusable protocol for comparing retrieval and selective abstention methods.
Motivation
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.