Search papers, labs, and topics across Lattice.
1
0
2
Memory-efficient optimizer state allocation can lead to dramatic improvements in training performance without sacrificing accuracy, as shown by SkewAdam's superior perplexity results.