TeleSWEBench
TeleSWEBench is a commit-driven benchmark with 734 questions derived from real developer commits in the srsRAN 5G repository. It evaluates LLM-powered software engineering agents in the telecommunications domain using executable unit tests and a hierarchical LLM-as-a-Judge framework across three difficulty tiers.
- Released
- 2026-06-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
General-purpose coding benchmarks fail to capture the stateful logic and strict requirements of telecom software, leaving a gap in evaluating ASE tools for this domain. TeleSWEBench provides a domain-specific benchmark with executable tests and a judge framework.
Motivation
With the telecommunications field embracing zero touch management alongside novel O-RAN and AI-RAN frameworks, contemporary telecom networks now function as immensely intricate and heavily softwareized codebases.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.