Search papers, labs, and topics across Lattice.
4
0
9
5
Self-evolving rubric rewards can dramatically enhance audio reasoning in models, outperforming traditional methods by adapting to the model's evolving capabilities.
Distilling from weaker models can enable a student to outperform its stronger counterparts, challenging the conventional wisdom of model hierarchy in AI training.
State-of-the-art LLMs fail to capture nuanced user preferences, lagging behind simple baselines in predicting choices in interactive narratives.
GFlowRL achieves unprecedented stability and performance in large language models by eliminating the problematic learned partition function, setting a new standard for GFlowNet-style reinforcement learning.