Search papers, labs, and topics across Lattice.
2
0
5
16
Kimi Delta Attention with Muon outperforms other architectures in validation loss, while Gated DeltaNet with AdamW leads in training throughput, revealing critical trade-offs in linear attention designs.
Training a 70B parameter open-source LLM on a supercomputer reveals the hidden engineering hurdles and infrastructure adaptations needed to democratize large-scale AI development beyond the private sector.