Search papers, labs, and topics across Lattice.
1
0
2
1
Kimi Delta Attention with Muon outperforms other architectures in validation loss, while Gated DeltaNet with AdamW leads in training throughput, revealing critical trade-offs in linear attention designs.