Search papers, labs, and topics across Lattice.
This study addresses the challenge of body-only emotion classification from skeleton motion in a difficult leave-performer-out evaluation setting. By employing an ensemble of eleven models with orthogonal error modes, the authors achieved a significant improvement in Macro-F1 score, reaching 36.80% compared to a baseline of 25.73%. Additionally, the research provides a tested explanation suite that reveals the importance of motion-grounded body-region evidence in decision-making, aligning well with Laban Movement Analysis attributes.
Combining models with orthogonal error modes leads to a 43% relative improvement in emotion recognition accuracy, revealing the critical role of body-region evidence in model decisions.
We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training performers, an equal-weight logit-mean ensemble reaches 36.80 +/- 4.00% per-fold Macro-F1, a protocol-matched +11.07 pp (+43% relative) over the same-split reproduced baseline. Our central contribution is a tested explanation suite: for a strong ensemble member, part-masking and counterfactual edits show (rather than assert) that its decisions depend on motion-grounded body-region evidence, and this region saliency aligns with rule-based Laban Movement Analysis (LMA) attributes far more than with classical kinematics: region-level saliency-LMA Spearman rho = +0.500 versus +0.033, roughly 15x, and the alignment holds for the submitted 11-way ensemble itself at rho = +0.517; the audit is post hoc and needs no retraining. The same suite faithfully reports a negative: within-window temporal saliency is diffuse rather than localized.