Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
2
Fine-grained MoE designs can outperform dense vision encoders while dramatically reducing latency, challenging the status quo in image and video understanding.
Current multi-modal LLMs struggle with the messy, real-world visual data captured by wearable devices, achieving only 24-52% accuracy on the new WearVQA benchmark.