Search papers, labs, and topics across Lattice.
This paper introduces PHILIA, a multi-robot agent designed to enhance long-term physical coexistence with intelligent robots by integrating a unified capability interface that decouples high-level reasoning from low-level execution. The architecture allows for a compositional user experience, enabling seamless upgrades across various components such as user interfaces and robot policies without requiring a complete system redesign. Validation on Astribot S1 robots demonstrates PHILIA's effectiveness in diverse household tasks, showcasing its ability to maintain contextual understanding and adapt to user preferences over extended interactions.
A unified capability interface in PHILIA allows for seamless upgrades in robot systems, enhancing user experience without the need for complete redesigns.
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstraction. PHILIA retains the rich interaction and tool ecosystem of OpenClaw while exposing robot-local runtimes, onboard perception, navigation, speaker, and robot policies through a unified capability interface. This design decouples low-frequency, high-semantic agent reasoning from high-frequency, low-level robot execution, enabling plug-and-play integration of user interfaces, robot embodiments, and policy backends. As a result, the user experience becomes compositional: advances in user interfaces, robot embodiments, robot policies, navigation, or interaction algorithms can improve the overall experience without redesigning the system. We validate the architecture on Astribot S1 robots while designing the robot gateway contract to support future heterogeneous robot platforms through a shared capability interface for observation, task execution, navigation, speech playback, status monitoring, and task cancellation. We present representative use cases in which agent memory and scene understanding are grounded in robot actions. These span interactive household scenarios, ranging from simple organization to challenging long-horizon and dexterous service tasks, such as packing a backpack and lifting a garbage bag. We highlight the human-robot interaction flow, where contextual understanding of user intent and preferences, together with human-in-the-loop confirmation or adjustment during execution, is essential for effective assistance.