Search papers, labs, and topics across Lattice.
University of Science and Technology
5
0
7
MLLMs can now seamlessly unify perception and reasoning, achieving state-of-the-art performance in free-form multimodal grounding tasks.
MoAKE achieves superior action quality assessment by unifying diverse action evaluations into a single model, overcoming the limitations of traditional one-by-one approaches.
GeoAnchor achieves superior 3D spatial reasoning by integrating multiple latent representations, outperforming existing methods that struggle with complex geometric scenarios.
SCoP-SAM not only sets a new benchmark for zinc-alloy intermediate phase segmentation but also excels in generalizing across various alloy scenarios.
Multi-modal models can now better handle distribution shifts thanks to a new method that explicitly models how different categories are distributed, even when the modalities are asymmetrical.