Search papers, labs, and topics across Lattice.
5
0
5
Mainstream video models consistently falter in fine-grained understanding, revealing critical vulnerabilities in their hallucination capabilities.
Long-horizon embodied agents struggle to translate long-term memory into actionable plans, exposing critical gaps in current benchmarks and methodologies.
AffordanceVLA transforms robotic manipulation by using structured affordance cues to create precise perception-action mappings, outperforming traditional models.
VLMs exhibit distinct failure modes under physical visual stress, revealing that traditional accuracy metrics can mask critical vulnerabilities in embodied AI systems.
Existing affordance prediction models fall flat when confronted with the wide-angle, distorted reality of panoramic vision, but a new training-free pipeline called PAP rises to the challenge.