Search papers, labs, and topics across Lattice.
This paper explores the use of vision foundation models for animal pose estimation and tracking, addressing the challenges posed by morphological diversity and limited annotated data. The authors propose two models鈥攐ne unsupervised and one supervised鈥攖hat leverage structural priors and diverse features to enhance tracking accuracy and cross-species robustness. Extensive evaluations on benchmarks like APTv2 and TigDog reveal that their approach strikes an effective balance between accuracy and generalization, making it a viable solution for wildlife monitoring and conservation efforts.
Vision foundation models can effectively track animal poses with minimal labeled data, achieving superior accuracy and cross-species robustness.
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data. Existing approaches either optimise generic keypoint localisation from annotated datasets (such as APTv2) with poor generalisation, or track custom keypoints using visual tracking, at the cost of performance. In this paper, we demonstrate that vision foundation models trained on large datasets can be used effectively to track animal pose with limited labelled data. We propose two models, one unsupervised and the other supervised, to track user-selected keypoints in videos. The supervised approach delivers superior tracking accuracy by employing a keypoint prompt encoder to explicitly inject structural priors from a reference frame into feature matching. In parallel, the unsupervised route provides strong cross-species robustness by leveraging diverse foundation-model features for training-free correspondence matching. Extensive evaluation on challenging animal video benchmarks APTv2 and TigDog demonstrates that our framework achieves strong performance while maintaining an effective balance between accuracy and generalisation, offering a practical solution for real-world animal behaviour analysis and conservation applications.