Benchmark Radar
AI BENCHMARK PROFILE

LibEvoBench

General AICoding & Software Engineering

LibEvoBench is a multi-task benchmark for evaluating code generation models on API evolution across versions of Python libraries. It includes tasks spanning multiple library versions and introduces the Software Evolution Understanding Score (SEUS) to measure version-specific knowledge consistency.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the gap in evaluating models on version-specific API knowledge, which is critical for real-world software maintenance. It highlights limitations of current training on temporally mixed corpora and motivates temporally grounded learning.

Motivation

Large software projects often depend on older versions of libraries, even as APIs continue to evolve across releases.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.