Benchmark Radar
AI BENCHMARK PROFILE

MirrorCode

General AICoding & Software EngineeringEpoch Research

MirrorCode evaluates AI agents on reimplementing entire software projects from behavior only, matching outputs on end-to-end tests across 25 programs spanning Unix utilities, data serialization, bioinformatics, interpreters, static analysis, cryptography, and compression.

Released
2026-06-29
Readiness
Runnable
Primary field
General AI

Why it matters

Existing coding benchmarks focus on shorter tasks, while long-horizon reimplementation remains hard to compare. MirrorCode provides a standardized, repeatable measure of autonomous software engineering capability.

Motivation

AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.