Benchmark Radar
AI BENCHMARK PROFILE

TensorBench

General AICoding & Software Engineering

TensorBench is a benchmark of 199 feature-addition and refactoring tasks on an open-source compiler-based tensor framework extending PyTorch. It grades agents by applying patches and running the framework's test suite.

Released
2026-06-04
Readiness
Paper only
Primary field
General AI

Why it matters

Repository-level coding benchmarks face a trade-off between difficulty and evaluation reliability. TensorBench uses automated test-based grading to provide reliable evaluation on challenging tasks.

Motivation

Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often involve large codebases with incomplete test coverage, while human review does not scale.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.