Benchmark Radar
AI BENCHMARK PROFILE

XREPOTEST

General AICoding & Software EngineeringSolis Team

Evaluates multilingual repository-level unit test generation across Rust, Go, Julia, PHP, and Ruby using containerized execution and context augmentation strategies.

Released
2026-08-26
Readiness
Runnable
Primary field
General AI

Why it matters

Exposes the gap between standalone and repository-level test generation, providing metrics like invocation rate to measure whether generated tests exercise intended functionality.

Motivation

Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.