Search papers, labs, and topics across Lattice.
This paper introduces a structured LLM architecture designed for zero-shot coordination between humans and robots in cooperative tasks with hidden goals, leveraging a Dec-POMDP framework. The approach integrates action-conditioned Theory-of-Mind inference, hierarchical planning, conversation interpretation, action verification, and feedback-based replanning to enhance decision-making. Experimental results show that this method outperforms both a baseline without ToM inference and a multi-agent reinforcement-learning policy, achieving fewer interaction steps and higher trust ratings from human participants.
Systematic decomposition of decision-making in human-robot coordination can significantly enhance trust and efficiency in collaborative tasks.
We present a structured large-language-model (LLM) architecture for zero-shot human--robot coordination in a cooperative construction task with private goal views. Guided by a Dec-POMDP formulation, the architecture decomposes decision-making into (i) action-conditioned Theory-of-Mind (ToM) inference, (ii) hierarchical planning, (iii) conversation interpretation, (iv) action verification, and (v) feedback-based replanning. We compare the proposed method with an ablation without ToM inference and a multi-agent reinforcement-learning policy trained offline over many goal pairs. In human-participant experiments, the proposed method required fewer interaction steps and yielded higher post-interaction trust ratings than both baselines. These results suggest that systematically decomposing the team decision problem, using LLMs as tractable surrogates for otherwise intractable inference and planning computations, and retaining conventional verification for physical feasibility can improve both task coordination and the human experience.