Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Even top-tier vision-language models consistently invert a subject's left and right, exposing a blind spot in camera- versus subject-centric spatial reasoning that rubric-guided GRPO over anatomical priors can finally resolve.
Ditching text-based chain-of-thought unlocks better audio-visual reasoning by interleaving textual steps with a unified latent space that preserves dense sensory information.