Benchmark Radar
AI BENCHMARK PROFILE

T2J-Bench

General AICoding & Software Engineering

T2J-Bench benchmarks codebase conversion by transferring PyTorch code to JAX under a fixed equivalence contract with three ordered verification stages.

Released
2026-05-27
Readiness
Paper only
Primary field
General AI

Why it matters

Reveals that agents overestimate success on codebase conversion, but the benchmark's data and verification harness are not publicly released.

Motivation

Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weakness: agents often over-trust their own local validation routines and declare success on artifacts that satisfy surface checks while violating the semantic contracts users actually care about.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.