Benchmark Radar
AI BENCHMARK PROFILE

RuBench

General AICoding & Software EngineeringEvgeny Shilov

Evaluates coding agents on 25 repository-level tasks in Russian, mined from recent fix commits across five open-source projects. Graded by upstream regression tests with withheld oracles. Multiple rounds document model change and contamination audits.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a benchmark with natively authored non-English specifications, addressing a gap in multilingual agent evaluation. Includes rigorous auditing and honest scores, which matter for reliability.

Motivation

Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.