Search papers, labs, and topics across Lattice.
The paper introduces HumanScore, a framework with six interpretable metrics to evaluate the quality of human motions in AI-generated videos, covering kinematic plausibility, temporal stability, and biomechanical consistency. Applying HumanScore to videos generated by thirteen state-of-the-art models reveals a gap between perceptual plausibility and motion biomechanical fidelity, highlighting failure modes like temporal jitter and anatomical implausibility. The framework provides robust model rankings based on quantitative and physically meaningful criteria.
AI-generated videos may look realistic, but HumanScore reveals they often fail at biomechanical fidelity, suffering from jitter, anatomically implausible poses, and motion drift.
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human bodies and motion dynamics. In this paper, we present HumanScore, a systematic framework to evaluate the quality of human motions in AI-generated videos. HumanScore defines six interpretable metrics spanning kinematic plausibility, temporal stability, and biomechanical consistency, enabling fine-grained diagnosis beyond visual realism alone. Through carefully designed prompts, we elicit a diverse set of movements at varying intensities and evaluate videos generated by thirteen state-of-the-art models. Our analysis reveals consistent gaps between perceptual plausibility and motion biomechanical fidelity, identifies recurrent failure modes (e.g., temporal jitter, anatomically implausible poses, and motion drift), and produces robust model rankings from quantitative and physically meaningful criteria.