Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Minibatch perturbations in AdamW can have delayed and significant impacts on future loss, reshaping our understanding of optimizer dynamics.
Simply adding more multimodal environments can hinder agent performance, but targeted diversity and structured difficulty can transform training outcomes.
SFT leads to task conflicts that can cripple multi-task learning, while RL's variance-limited updates enable seamless task coexistence.