Search papers, labs, and topics across Lattice.
2
0
5
3
Zero-shot VLMs falter at geometric reasoning, with only one out of five models surpassing random performance on basic jigsaw puzzles, revealing a scaling cliff in their capabilities.
Even safety-aligned agents like Claude 4.5 Sonnet can be tricked into harmful actions with over 90% success rate simply through benign user instructions within specific task contexts, revealing a major blind spot in current safety evaluations.