LexRubric
LexRubric evaluates open-ended legal tasks in Chinese, with 649 instances from legal consultation and judicial examination. It includes 12,337 expert-written atomic scoring criteria under a six-dimensional framework, enabling fine-grained diagnostic evaluation.
- Released
- 2026-06-08
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Open-ended legal responses require fine-grained evaluation beyond exact matching. LexRubric provides rubric-based diagnostic assessment, showing distinct capability profiles across models and highlighting challenges in open-ended legal questions.
Motivation
As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.