Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
10
Fine-grained MoE designs can outperform dense vision encoders while dramatically reducing latency, challenging the status quo in image and video understanding.
Open-ended generation models are failing to capture factual completeness, with top performers only achieving 58.7% accuracy on a new benchmark designed to assess this critical aspect.
Current multi-modal LLMs struggle with the messy, real-world visual data captured by wearable devices, achieving only 24-52% accuracy on the new WearVQA benchmark.