Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Current vision-language models excel at recognizing objects but falter in capturing dynamic interactions and user intent over time, revealing critical gaps in embodied AI.
Explicitly aligning hand actions based on temporal offsets boosts segmentation accuracy by over 2 F1 points with minimal additional parameters.