Benchmark Radar
AI BENCHMARK PROFILE

WuYu-EnvLE-Bench

General AIKnowledge & Reasoning

WuYu-EnvLE-Bench evaluates LLMs on environmental law enforcement with 2,521 instances across 14 tasks and 12 pollution-medium subdomains, covering pre-, in-, and post-enforcement workflows. It uses Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI) for evaluation.

Released
2026-07-20
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark fills a gap in evaluating LLMs for evidence-grounded, rule-aware reasoning in environmental enforcement. It provides practical assessment of model reliability in traceable decision-making, highlighting limitations in evidence-chain construction and procedural judgment, which is critical for legal and regulatory applications.

Motivation

Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.