Benchmark Radar
AI BENCHMARK PROFILE

TeleSWEBench

General AICoding & Software Engineering

TeleSWEBench is a commit-driven benchmark with 734 questions derived from real developer commits in the srsRAN 5G repository. It evaluates LLM-powered software engineering agents in the telecommunications domain using executable unit tests and a hierarchical LLM-as-a-Judge framework across three difficulty tiers.

Released
2026-06-03
Readiness
Paper only
Primary field
General AI

Why it matters

General-purpose coding benchmarks fail to capture the stateful logic and strict requirements of telecom software, leaving a gap in evaluating ASE tools for this domain. TeleSWEBench provides a domain-specific benchmark with executable tests and a judge framework.

Motivation

With the telecommunications field embracing zero touch management alongside novel O-RAN and AI-RAN frameworks, contemporary telecom networks now function as immensely intricate and heavily softwareized codebases.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.