Search papers, labs, and topics across Lattice.
Peking University
5
0
8
8
CODA redefines edge video diffusion by achieving up to 1.80x faster inference and 1.74x better energy efficiency without sacrificing output quality.
Filtering out high-frequency tokens from LLM embeddings can significantly enhance their semantic quality and downstream performance.
Current image restoration models still fail to strike the right balance between noise reduction, detail fidelity, and accurate color in real-world, low-light portrait scenarios, highlighting a critical gap this challenge aims to close.
Open-source ATLAS unlocks rapid, accurate co-design of 3D-DRAM memory systems and LLM accelerators, previously hindered by closed-source tools and customized designs.
A hybrid-bonding-based LLM serving accelerator, Helios, tackles the dynamic nature of KV cache management in LLM serving, achieving significant speedup and energy efficiency gains over existing GPU/NMP designs.