Search papers, labs, and topics across Lattice.
5
0
6
24
Achieving a 28% improvement in alignment performance with just 100 preference samples highlights the potential of meta-learning to bridge the data gap in multilingual LLMs.
Routing decisions based on exploratory trajectories can significantly boost cost efficiency in software engineering tasks without losing the performance edge of stronger models.
Confidence-based remasking in dLLMs may not deliver the expected improvements and can actually worsen diversity issues in certain decoding settings.
Ditch the ELBO: bypassing biased likelihood approximations in RL fine-tuning of diffusion LMs unlocks more stable and effective policy optimization, yielding nearly 20% accuracy gains on challenging tasks.
A 3B model, guided by a novel RL framework, can outperform a 20B model in capturing diverse human perspectives, challenging the assumption that larger models inherently possess better alignment.