Search papers, labs, and topics across Lattice.
5
0
9
6
RxBrain achieves a groundbreaking integration of language and visual reasoning, enabling agents to generate embodied plans that seamlessly connect abstract tasks with physical actions.
Terminal-use agents are still far from achieving reliable general-purpose performance, with top models only scoring 65.8% on a new benchmark that spans diverse real-world tasks.
Ditching the vision encoder actually *improves* multimodal understanding at scale, proving that pixel embeddings alone can achieve state-of-the-art results in unified multimodal models.
LALMs struggle more with *hearing* the evidence than *reasoning* about it, and EvA's evidence-first fusion architecture proves it.
Robots can now recover from failures during manipulation tasks by explicitly tracking progress against spatial subgoals, without needing extra training data or models.