TensorBench
TensorBench is a benchmark of 199 feature-addition and refactoring tasks on an open-source compiler-based tensor framework extending PyTorch. It grades agents by applying patches and running the framework's test suite.
- Released
- 2026-06-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Repository-level coding benchmarks face a trade-off between difficulty and evaluation reliability. TensorBench uses automated test-based grading to provide reliable evaluation on challenging tasks.
Motivation
Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often involve large codebases with incomplete test coverage, while human review does not scale.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.