AI BENCHMARK PROFILE
T2J-Bench
T2J-Bench benchmarks codebase conversion by transferring PyTorch code to JAX under a fixed equivalence contract with three ordered verification stages.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Reveals that agents overestimate success on codebase conversion, but the benchmark's data and verification harness are not publicly released.
Motivation
Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weakness: agents often over-trust their own local validation routines and declare success on artifacts that satisfy surface checks while violating the semantic contracts users actually care about.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.