AI BENCHMARK PROFILE
Kotlin Benchmark
Evaluates AI coding agents on real-world Kotlin tasks from nine open-source repositories. The benchmark uses reproducible Docker environments, regression tests, and a scoring protocol that awards pass only when all expected test transitions are met.
- Released
- 2026-07-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
This benchmark addresses the lack of reference evaluation for Kotlin-specific coding agents, providing a reproducible, task-level framework for comparing agent performance on real-world Kotlin issues and tracking progress.
Motivation
Measure whether general coding agents can solve realistic Kotlin repository issues under a reproducible SWE-bench-style protocol.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.