SA-Bench
SA-Bench (SemanticAlign-Bench) evaluates semantic alignment in LLM-based paper reproduction across 30 papers from top conferences. It decomposes paper specifications into atomic verifiable claims (SAUs) and evaluates repositories along four diagnostic dimensions (numerical, methodological, protocol, ordering drift). Includes 1,491 SAUs across five ML domains and evaluates 12 generator configurations.
- Released
- 2026-08-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LLM agents generating code for paper reproduction often produce semantically unfaithful implementations. SA-Bench provides a diagnostic framework to measure semantic drift, revealing that current agents struggle with faithful implementation, guiding development of better scaffolding.
Motivation
LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.