Search papers, labs, and topics across Lattice.
This paper introduces RuleWorld, a benchmark designed to evaluate large language models' (LLMs) ability to understand and apply procedural rules in various reasoning scenarios. It also presents DynaRule, an innovative framework that enhances LLMs by integrating rules into the KV cache and employing Stacked Step-Level Attention Training to facilitate dynamic rule re-attention during inference. Experimental results demonstrate that DynaRule significantly improves QA accuracy by up to 19 points and achieves over 85% Recall@1 with 10K rules, outperforming existing models in complex reasoning tasks.
LLMs can now dynamically re-attend to procedural rules, boosting their reasoning accuracy by up to 19 points in complex scenarios.
Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large-scale benchmark that reformulates rules as globally reusable abstract units rather than instance-specific facts. In RuleWorld, several scenarios, including single-rule, parallel multi-rule, and multi-hop reasoning, are settled for comprehensive evaluation. We further propose DynaRule, an end-to-end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step-wise process. Specifically, DynaRule employs Stacked Step-Level Attention Training with a special <search> token to enable dynamic rule re-attention and updating during inference. In this way, the model can re-attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi-step reasoning. Experiments on RuleWorld show that existing LLMs face challenges under large rule pools, while DynaRule improves average QA accuracy by up to 19 points and achieves over 85% Recall@1 at 10K rules, outperforming strong baselines by large margins. We make our code and dataset available here: https://github.com/SharkSpicy-NLP/Beyond-Factual-Knowledge.