Search papers, labs, and topics across Lattice.
University of Maryland, College Park
4
0
7
2
Adjusting the KL penalty in self-distillation can transform a brittle method into a robust framework that significantly boosts reasoning performance.
Optimizing the order of thought in diffusion models can boost accuracy by over 9% on complex tasks like Sudoku and mathematical reasoning.
LLMs exhibit a pervasive optimism bias when evaluating research proposals, frequently rating methodologically unsound ideas as promising, suggesting they're not ready to replace human reviewers.
Instead of imitating reflections, LLM agents can be trained to reason about action quality by rewarding correct judgments between alternative actions, leading to improved performance and generalization.