Search papers, labs, and topics across Lattice.
This paper introduces a physics-consistent benchmark for evaluating contact-rich human-robot interactions, specifically in the context of robot-assisted bathing. The benchmark incorporates a physically responsive human model and physics-aware scoring to assess interaction quality beyond mere task completion. Results reveal that while an LLM-augmented state machine achieves a 72.9% task success rate, significant drops occur when applying physical validity checks, underscoring the inadequacy of traditional evaluation methods in capturing real-world interaction failures.
Task completion rates can be misleading; a physics-aware evaluation reveals that many robotic interactions fail to ensure safe and effective contact with humans.
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss failures that emerge only during physical human contact. This limitation is critical in contact-rich assistive tasks, where meaningful evaluation requires a physically responsive human, interaction-quality assessment beyond task success, and a leak-free observer-scorer protocol. We introduce a physics-consistent benchmark for contact-rich human-robot interaction, instantiated in robot-assisted bathing. The benchmark combines a deformable, passively responding human, physics-aware scores alongside task-level success, and a frozen vision-only / scorer-only evaluation protocol. To establish physical validity, region-wise simulated responses are calibrated against force-indentation measurements from Franka impedance pushes on a medical-care manikin. Under a frozen T1-T7 protocol with 140 runs per method, an LLM-augmented state machine (State Machine) achieves 72.9% task success but drops to 56.4% after correct-region and force-safety screening; VoxPoser produces lighter and more stable contact but completes only 27.9% of trials; and zero-shot pi0.5 achieves 0.7% task success with no correct-region or safety-gated successes. These results show that task completion alone does not imply physically valid contact and motivate physics-aware screening before deployment of contact-rich assistive robot policies.