Benchmark Radar
AI BENCHMARK PROFILE

DDBench

General AICoding & Software Engineering

DDBench is a code-repair benchmark with 60 historical bugs from 13 open-source distributed systems, evaluated under symptom-only and context-augmented conditions.

Released
2026-08-14
Readiness
Paper only
Primary field
General AI

Why it matters

It isolates the effect of debugging context on agent success in distributed-system repair, addressing a gap left by single-process benchmarks.

Motivation

LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.