AI BENCHMARK PROFILE
XREPOTEST
Evaluates multilingual repository-level unit test generation across Rust, Go, Julia, PHP, and Ruby using containerized execution and context augmentation strategies.
- Released
- 2026-08-26
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Exposes the gap between standalone and repository-level test generation, providing metrics like invocation rate to measure whether generated tests exercise intended functionality.
Motivation
Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.