Search papers, labs, and topics across Lattice.
2
0
5
5
It is shown that this gating mechanism causes SSMs to first learn an in-weights"memorization"solution, while delaying, or even preventing, convergence to a correct in-context learning solution, and that gating is often beneficial for improving generalization to long sequence lengths.
Tri-modal masked diffusion models can now be trained from scratch, achieving strong results in text generation, text-to-image, and text-to-speech, thanks to a systematic exploration of the design space and a novel SDE-based batch size reparameterization.