Search papers, labs, and topics across Lattice.
Bytedance Seed
2
0
3
Fast-weight attention can enhance language modeling and improve task performance by effectively managing context and memory in recurrent networks.
MultiMDM achieves superior few-step generation by cleverly preserving the masking structure, enabling a drafting capability that enhances token refinement.