Search papers, labs, and topics across Lattice.
This paper introduces DIRECT, a novel framework designed to enhance sequence labeling tasks by improving domain alignment and inference efficiency in large language models. By employing Direct Preference Optimization (DPO) post-supervised fine-tuning and a controlled decoding process, DIRECT ensures that outputs adhere to specified formats and candidate sets. Experimental results across eight datasets reveal that DIRECT significantly outperforms existing methods in both accuracy and computational efficiency.
DIRECT achieves remarkable efficiency gains in sequence labeling by leveraging a template-filling mechanism that minimizes redundant computations while enhancing task alignment with human preferences.
Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alignment and low inference efficiency. To address these issues, we propose DIRECT, a framework that addresses these issues through training-time optimization and inference-time rectification. Specifically, DIRECT performs Direct Preference Optimization (DPO) after supervised fine-tuning to strengthen task alignment with human preferences, and introduces a controlled decoding process that enforces fixed output formats and restricts predictions to candidate sets. To further improve efficiency, a template-filling mechanism requires the model to generate only label tokens while reusing prefixed content through the KV Cache, thus reducing redundant computation. Experimental results on eight datasets demonstrate that DIRECT achieves significant improvements in both performance and efficiency compared to existing methods.