LibEvoBench
LibEvoBench is a multi-task benchmark for evaluating code generation models on API evolution across versions of Python libraries. It includes tasks spanning multiple library versions and introduces the Software Evolution Understanding Score (SEUS) to measure version-specific knowledge consistency.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating models on version-specific API knowledge, which is critical for real-world software maintenance. It highlights limitations of current training on temporally mixed corpora and motivates temporally grounded learning.
Motivation
Large software projects often depend on older versions of libraries, even as APIs continue to evolve across releases.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.