Search papers, labs, and topics across Lattice.
10
0
11
2
The 9B WeMM-Embedding variant sets a new benchmark in multimodal embeddings, outperforming larger models while being deployed at scale across WeChat's diverse applications.
Anchoring positional embeddings in video generation can drastically enhance the quality of looping videos, achieving unprecedented temporal coherence and visual fidelity.
PaDoc achieves a remarkable 94.24 F1 score among end-to-end parsers while being the fastest parser at five concurrency levels, revolutionizing document parsing efficiency.
Orca's unified world latent space enables superior performance in diverse tasks, outperforming specialized models with a single framework.
Integrating real images into the GRPO process and using dual-reward guidance allows PortraitGen to achieve unprecedented levels of photorealism while effectively suppressing AI artifacts.
WATER-S is a game-changing dataset that boosts WordArt recognition capabilities by hundreds of times, enabling unprecedented accuracy in complex text layouts.
WeGenBench exposes the hidden deficiencies of text-to-image models, revealing that many leading systems struggle with specific generation tasks despite overall high performance.
Finally, a unified framework lets you control both facial appearance and voice timbre for personalized audio-video generation across multiple identities.
LLM agents can now leverage a unified memory framework that dynamically adapts to different question types, enabling more coherent and user-centric long-horizon dialogues.
Cycle consistency unlocks SOTA cross-view object correspondence in videos without ground-truth annotations, even enabling test-time training.