Search papers, labs, and topics across Lattice.
Xidian University
2
0
4
OTF-CBM reveals that dynamic transport processes can significantly enhance the interpretability and accuracy of vision-language models, outperforming static alignment methods.
By injecting LLM-derived contextual cues into skeleton representations, SkeletonContext achieves state-of-the-art zero-shot action recognition, even distinguishing visually similar actions without explicit object interactions.