Search papers, labs, and topics across Lattice.
The paper introduces IRWOZ 2.0, an enhanced dialogue dataset for industrial human-robot interaction that leverages large language models (LLMs) to improve the quality and accuracy of dialogue states and utterances. By addressing the noise found in the original IRWOZ dataset, the authors expand the dataset to 390 dialogues across four industrial domains and implement manual corrections alongside automated typo removal. Benchmark experiments reveal a substantial increase in dialogue state tracking performance, with BLEU-4 scores improving from 0.1651 to 0.5604, underscoring the dataset's utility for advancing HRI research.
Dialogue state tracking accuracy in industrial robots skyrockets with the new IRWOZ 2.0 dataset, achieving a BLEU-4 score increase of over 200%.
IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2's BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ. To support industrial HRI research, we publicly released IRWOZ 2.0 dataset at https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations