Search papers, labs, and topics across Lattice.
This paper introduces a cloud-edge multimodal interaction framework designed for robots, enhancing gesture perception and task planning in complex environments. By integrating an improved YOLO-based gesture detector with large language and vision-language models, the system achieves high precision in gesture detection and effective action planning despite limited onboard computing resources. Experimental results indicate that the system can successfully execute various tasks with high satisfaction rates, showcasing its potential for robust human-robot interaction.
Achieving 98.9% precision in gesture detection, this framework redefines how robots can interact in resource-constrained environments.
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.