Benchmark Radar
AI BENCHMARK PROFILE

LogDx-CI

General AICoding & Software EngineeringLogDx-CI Team

Benchmark evaluating 11 log reduction tools on 35 GitHub Actions failure cases, scored by 3 LLM debugger families.

Released
2026-05-26
Readiness
Paper only
Primary field
General AI

Why it matters

No public comparison existed for which log reductions preserve diagnostic evidence for LLM-based root-cause diagnosis.

Motivation

CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.