Search papers, labs, and topics across Lattice.
Kling Team, Kuaishou Technology
4
0
5
CineCap achieves a new state of the art in cinematographic video captioning by effectively balancing descriptive completeness with factual accuracy through innovative structured reasoning techniques.
Text-to-image customization can now preserve the original model's behavior, thanks to a decoupled learning objective that balances new concepts with pre-existing capabilities.
Forget single-number video quality scores: UltraVQA and Analytic Score Optimization (ASO) unlock richer, multi-faceted evaluations that better align with human preferences.
Forget generic CoT: Embed-RL uses reinforcement learning to generate reasoning traces that are explicitly optimized for multimodal embedding tasks, leading to significant performance gains.