Search papers, labs, and topics across Lattice.
2
0
4
2
Current MLLMs struggle with active visual observation, scoring as low as 3.5% on tasks designed to test this critical cognitive function.
Text-to-image models wash away the unique stylistic fingerprints of their captioning counterparts, revealing a surprising disconnect between text and image generation.