Benchmark Radar
AI BENCHMARK PROFILE

RepoReasoner

General AICoding & Software EngineeringLong Context & MemoryRepoReasoner Team

Evaluates long-context LLMs on repository-level code reasoning through Output Prediction and Call Chain Prediction tasks, using dynamic tracing and I/O rewriting to reduce memorization.

Released
2026-07-28
Readiness
Paper only
Primary field
General AI

Why it matters

Assesses cross-file reasoning capabilities that are critical for real-world software engineering, identifying limitations beyond function-level benchmarks.

Motivation

Recent large language models (LLMs) have shown strong performance on software engineering tasks, yet most existing benchmarks evaluate code reasoning at the function level, where all relevant information is localized.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.