Search papers, labs, and topics across Lattice.
WeChat Vision, Tencent Inc.
4
0
8
1
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
Achieving up to 46.1% lower latency and 67.4% reduced data transmission, this framework redefines efficient and privacy-aware LLM inference across diverse devices.
Thinking Collapse can severely impair reasoning in LLMs, but a new adaptive framework boosts accuracy by over 4% while preserving cognitive capacity.
Removing objects from video now means removing their shadows and reflections too, thanks to a new method that teaches diffusion models to "understand" object-scene physics.