Search papers, labs, and topics across Lattice.
This paper introduces MistyPilot, a multi-agent framework that enables the control of social robots through natural-language instructions by orchestrating various skills. By employing a Task Router to allocate tasks between a Physically Interactive Agent and a Social Interaction Agent, the system efficiently manages both physical interactions and dialogue-driven task states. Evaluation results show high accuracy in skill execution and positive user feedback, indicating that MistyPilot significantly enhances the usability of social robots in interactive environments.
MistyPilot achieves high accuracy in multi-agent orchestration, enabling social robots to seamlessly execute complex tasks from natural language instructions.
Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often require developers to manually compose APIs into skills, configure their parameters, bind sensor events to skills, and manage task states at runtime. We present MistyPilot, a multi-agent LLM framework that interprets high-level natural-language instructions and orchestrates the corresponding skills on the Misty social robot. A Task Router dispatches each instruction to one of two specialized agents: a Physically Interactive Agent for sensor-triggered robot control and direct skill invocation, and a Social Interaction Agent for dialogue-oriented task-state management and context-dependent multimodal response generation. To improve efficiency, the Social Interaction Agent reuses previously generated results when applicable and invokes full generation otherwise. We evaluate MistyPilot on five component-level suites, with sensor bindings and skill invocations executed on the physical Misty robot, and a preliminary user study with 12 participants. MistyPilot attains high accuracy on routing, sensor-skill binding, task-state parsing, result reuse, and skill extension up to 100 skills, and lower variance than an otherwise identical single-agent baseline, while participants report positive perceptions of usability and interaction quality. The code will be made publicly available via the project page.