Benchmark Radar
AI BENCHMARK PROFILE

Multi-LCB

General AICoding & Software EngineeringMulti-LCB Team

Multi-LCB evaluates code generation across twelve programming languages (C++, C#, Python, Java, Rust, Go, TypeScript, JavaScript, Ruby, Kotlin, Scala, PHP) by transforming Python tasks from LiveCodeBench while preserving its contamination controls and evaluation protocol.

Released
2026-06-18
Readiness
Runnable
Primary field
General AI

Why it matters

LiveCodeBench restricted code evaluation to Python; Multi-LCB addresses the gap by enabling cross-language assessment, revealing Python overfitting and language-specific contamination in LLMs, and supporting robust multilingual code evaluation for real-world software engineering.

Motivation

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.