Search papers, labs, and topics across Lattice.
Xi'an Jiaotong University
2
0
3
IntentQA reveals that understanding intent in videos requires more than just visual recognition, highlighting the critical role of cognitive context in achieving robust performance.
Geo-Embed achieves a 15.3% performance boost over existing models, redefining how we approach multimodal urban understanding.