Benchmark Radar
AI BENCHMARK PROFILE

Multi-Legal-Bench

General AICoding & Software Engineering

Multi-Legal-Bench evaluates LLMs on legal reasoning across six countries, four language families, and five tasks, using structured metadata from court registries for classification and extraction.

Released
2026-05-28
Readiness
Inspectable
Primary field
General AI

Why it matters

It enables cross-lingual and cross-jurisdictional comparison of legal reasoning, revealing that transfer quality depends more on label-set alignment than language proximity, aiding model selection in legal NLP.

Motivation

Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.