Benchmark Radar
AI BENCHMARK PROFILE

UOJ-Bench

General AICoding & Software Engineering

UOJ-Bench is a benchmark for evaluating LLMs in code generation, hacking, and repair, built from real-world submissions on the Universal Online Judge and evaluated through UOJ's judging infrastructure.

Released
2026-06-11
Readiness
Paper only
Primary field
General AI

Why it matters

UOJ-Bench extends beyond problem-solving to include identifying errors in human code, a critical educational activity. It provides a realistic setting for assessing LLM support in learning.

Motivation

Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.