AI BENCHMARK PROFILE
Multi-Legal-Bench
Multi-Legal-Bench evaluates LLMs on legal reasoning across six countries, four language families, and five tasks, using structured metadata from court registries for classification and extraction.
- Released
- 2026-05-28
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
It enables cross-lingual and cross-jurisdictional comparison of legal reasoning, revealing that transfer quality depends more on label-set alignment than language proximity, aiding model selection in legal NLP.
Motivation
Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.