Search papers, labs, and topics across Lattice.
This paper introduces WuYu-EnvLE-Bench, a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in the context of environmental law enforcement, comprising 2,521 instances across various tasks and pollution domains. The evaluation reveals that while LLMs excel in structured, rule-based tasks, they struggle significantly with complex reasoning aspects such as evidence-chain construction and contradiction detection. Notably, the findings indicate that scaling models does not necessarily lead to improved performance in these critical reasoning areas, underscoring the need for more sophisticated approaches to enforcement reasoning in AI applications.
LLMs may ace rule-based tasks, but they falter in crucial reasoning areas like evidence integration and contradiction detection, revealing a significant gap in their utility for environmental law enforcement.
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.