WuYu-EnvLE-Bench
WuYu-EnvLE-Bench evaluates LLMs on environmental law enforcement with 2,521 instances across 14 tasks and 12 pollution-medium subdomains, covering pre-, in-, and post-enforcement workflows. It uses Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI) for evaluation.
- Released
- 2026-07-20
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
The benchmark fills a gap in evaluating LLMs for evidence-grounded, rule-aware reasoning in environmental enforcement. It provides practical assessment of model reliability in traceable decision-making, highlighting limitations in evidence-chain construction and procedural judgment, which is critical for legal and regulatory applications.
Motivation
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.