AI BENCHMARK PROFILE
DDBench
DDBench is a code-repair benchmark with 60 historical bugs from 13 open-source distributed systems, evaluated under symptom-only and context-augmented conditions.
- Released
- 2026-08-14
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It isolates the effect of debugging context on agent success in distributed-system repair, addressing a gap left by single-process benchmarks.
Motivation
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.