Search papers, labs, and topics across Lattice.
4
0
6
0
Instead of trusting black-box scalar scores to evaluate world models, video verification can match end-to-end VLM recall while producing fully auditable, spatiotemporally grounded physical proofs of failure.
Long-tail semantic failures in document understanding are exposed when models are forced to reason with a unified vocabulary of visual anchors rather than treating elements in isolation.
Achieving over 134% accuracy gains with just 1,000 trainable parameters reveals a game-changing approach to enhancing spatial reasoning in vision language models.
SeClaw reveals that existing benchmarks fall short in capturing the complexities of agent behavior, enabling a more nuanced evaluation of security risks in autonomous systems.