Search papers, labs, and topics across Lattice.
9
0
11
9
Mid-training with function-aware fill-in-the-middle boosts coding agent performance while preventing capability erosion in non-agentic tasks.
Generators can dramatically improve their performance on long-tailed visual requests by leveraging a teach-then-search co-training approach, overcoming a critical knowledge boundary.
DR-DCI achieves a remarkable 73.3% accuracy in agentic search tasks while efficiently scaling from 100K to 10M documents, outperforming traditional methods.
Even the top-performing MLLMs struggle with visual reasoning, achieving only 64% accuracy on a benchmark designed to reflect real-world diversity.
MiniMax-M2 proves that massive parameter counts don't always translate to better agentic performance; strategic activation of a smaller subset can unlock frontier-level intelligence.
Today's visual generation models are often evaluated on the wrong things, leading to inflated performance claims that mask critical failures in spatial reasoning, temporal consistency, and causal understanding.
Today's best AI agents can only complete 33% of common online tasks like booking appointments or filling out job applications, revealing a significant gap between current capabilities and real-world utility.
Image generation models ace photorealistic art but still choke on screenshots and infographics, highlighting a critical gap in real-world applicability.
MLLMs may ace your visual question answering, but VisPhyWorld reveals they're still struggling to actually *simulate* physics.