Search papers, labs, and topics across Lattice.
3
0
6
0
Language can effectively guide pixel-level anomaly detection without compromising visual fidelity, leading to unprecedented performance in industrial applications.
Current MLLMs are still surprisingly reliant on textual reasoning, even when visual information is crucial for solving STEM problems.
LongCat-Next shatters the language-centric paradigm by unifying text, vision, and audio into a single autoregressive model with minimal modality-specific design, finally reconciling understanding and generation in discrete vision modeling.