AI BENCHMARK PROFILE
LogDx-CI
Benchmark evaluating 11 log reduction tools on 35 GitHub Actions failure cases, scored by 3 LLM debugger families.
- Released
- 2026-05-26
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
No public comparison existed for which log reductions preserve diagnostic evidence for LLM-based root-cause diagnosis.
Motivation
CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.