Search papers, labs, and topics across Lattice.
4
0
6
9
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Current audio-visual generation models struggle to maintain coherence and alignment when scaling to minute-long content, a problem exposed by the new LongAV-Compass benchmark.
Stop reinventing the wheel: OpenWorldLib offers a unified framework and codebase for advanced world models, finally bringing standardization to a fragmented field.