Search papers, labs, and topics across Lattice.
6
0
6
5
Achieving 74.8% accuracy on a new temporal reasoning benchmark, ChronoVision redefines how multimodal models can tackle complex visual tasks.
GraphVid achieves superior video quality and controllability with significantly less training data, revolutionizing how we can interact with multi-object video generation.
ELSA3D outperforms existing unified 3D models by halving computational costs while enhancing cross-modal reasoning precision.
Forget brute-force VLMs: parameter-efficient fine-tuning with high-quality rationales unlocks surprisingly accurate and interpretable time-series anomaly detection.
Training video generation models to explicitly infer latent physical properties yields more physically plausible videos than simply scaling data and model size.
Text-to-3D generation gets a semantic upgrade: DreamPartGen creates 3D objects with parts that not only look right but also understand their relationships and align with textual descriptions.