Benchmark Radar
AI BENCHMARK PROFILE

RepoMirage

General AICoding & Software Engineering

RepoMirage is a two-stage evaluation suite built on SWE-Bench Verified that applies semantics-preserving repository-level perturbations and extended tasks to probe repository context reasoning in code agents.

Released
2026-05-25
Readiness
Paper only
Primary field
General AI

Why it matters

It aims to isolate repository context reasoning from end-to-end issue resolution performance, revealing a gap that could inform structure-aware agent design.

Motivation

Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on end-to-end tasks such as issue resolution truly reflects repository context reasoning, the ability to identify the task-relevant information across multiple files and reason over the relations among them.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.