Benchmark Radar
AI BENCHMARK PROFILE

SA-Bench

General AIKnowledge & Reasoning

SA-Bench (SemanticAlign-Bench) evaluates semantic alignment in LLM-based paper reproduction across 30 papers from top conferences. It decomposes paper specifications into atomic verifiable claims (SAUs) and evaluates repositories along four diagnostic dimensions (numerical, methodological, protocol, ordering drift). Includes 1,491 SAUs across five ML domains and evaluates 12 generator configurations.

Released
2026-08-25
Readiness
Paper only
Primary field
General AI

Why it matters

LLM agents generating code for paper reproduction often produce semantically unfaithful implementations. SA-Bench provides a diagnostic framework to measure semantic drift, revealing that current agents struggle with faithful implementation, guiding development of better scaffolding.

Motivation

LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.