Search papers, labs, and topics across Lattice.
I,⋯,oi,G
4
0
6
2
MemTrain reveals that self-supervised memory training can outperform traditional reinforcement learning approaches in enhancing LLMs' reasoning capabilities.
TrOPD stabilizes on-policy distillation by ensuring reliable teacher supervision, leading to consistent performance improvements over existing methods.
Stop cobbling together memory-augmented agents: MemFactory offers a unified "Lego-like" framework that streamlines training and boosts performance by up to 14.8%.
Pointwise reward models can finally compete with pairwise models in RLHF, thanks to a new intergroup comparison method that scales linearly with the number of candidates.