AI BENCHMARK PROFILE
UOJ-Bench
UOJ-Bench is a benchmark for evaluating LLMs in code generation, hacking, and repair, built from real-world submissions on the Universal Online Judge and evaluated through UOJ's judging infrastructure.
- Released
- 2026-06-11
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
UOJ-Bench extends beyond problem-solving to include identifying errors in human code, a critical educational activity. It provides a realistic setting for assessing LLM support in learning.
Motivation
Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.