RuBench
Evaluates coding agents on 25 repository-level tasks in Russian, mined from recent fix commits across five open-source projects. Graded by upstream regression tests with withheld oracles. Multiple rounds document model change and contamination audits.
- Released
- 2026-07-07
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a benchmark with natively authored non-English specifications, addressing a gap in multilingual agent evaluation. Includes rigorous auditing and honest scores, which matter for reliability.
Motivation
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.