RepoMirage
RepoMirage is a two-stage evaluation suite built on SWE-Bench Verified that applies semantics-preserving repository-level perturbations and extended tasks to probe repository context reasoning in code agents.
- Released
- 2026-05-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It aims to isolate repository context reasoning from end-to-end issue resolution performance, revealing a gap that could inform structure-aware agent design.
Motivation
Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on end-to-end tasks such as issue resolution truly reflects repository context reasoning, the ability to identify the task-relevant information across multiple files and reason over the relations among them.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.