Search papers, labs, and topics across Lattice.
Monash University, Shanghai University of Electric Power
7
0
8
Current spatio-temporal video grounding models falter on complex queries, with performance dropping sharply when faced with compositional challenges.
Achieving leading performance in image generation with only 6 billion parameters, Swift-Image redefines the efficiency frontier for compact models.
A novel capability-driven data infrastructure enables multimodal models to achieve unprecedented versatility and transferability in image generation tasks.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
Causal attention in Post-Norm Transformers amplifies token similarity, leading to a collapse that training dynamics fail to repair, revealing critical insights into model behavior.
Finally, a virtual try-on system that can handle extreme poses, lighting variations, and motion blur while preserving garment texture and material properties in near real-time.
Forget trajectory-level rollouts: MuSEAgent learns faster and reasons better by distilling past interactions into reusable, state-aware decision experiences.