Search papers, labs, and topics across Lattice.
7
0
11
11
Memory management emerges as a high-leverage skill that can double or quadruple the performance of LLMs in complex tasks without altering their core action behaviors.
Learning from failures can boost agent success rates by over 6% without extra training, reshaping how we approach agent improvement.
Jointly training MTP and RL doesn't have to hurt: a simple coefficient calibration scheme unlocks performance gains on mathematical reasoning tasks.
LLM unlearning via counterfactual tuning can backfire, increasing hallucination rates in unexpected areas due to inconsistencies in the "fake" knowledge it's trained on.
Even GPT-4 struggles to maintain its reasoning prowess when faced with the rigor and efficiency demands of a realistic high school exam, suggesting current LMMs are far from being ready for prime time as intelligent tutors.
Overcome the prohibitive cost of ground-truth labels in reinforcement learning by actively acquiring labels for only the most valuable samples, leading to stable training and improved performance even with limited annotation budgets.
ELBO-based reinforcement learning, previously dismissed for visual generation, can actually outperform MDP-based methods for aligning denoising generative models with human preferences.