Benchmark Radar
AI BENCHMARK PROFILE

Doc2CI

General AICoding & Software Engineering

DOC2CI evaluates LLM-generated CI/CD configuration YAML against reference configurations from four CI services, measuring exact match and schema validity.

Released
2026-08-02
Readiness
Paper only
Primary field
General AI

Why it matters

It quantifies the gap between LLM-generated configuration validity and reference similarity, highlighting the need for schema-aware evaluation in configuration generation.

Motivation

Adopting Continuous Integration (CI) often requires writing YAML configurations that are error-prone and challenging to maintain.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.