Search papers, labs, and topics across Lattice.
5
1
8
3
Early persona integration in language models can drastically reduce misalignment in moral dilemmas and enhance adherence to desired values.
Base models outperform post-trained models in emulating human opinions, revealing a crucial distinction in how we should approach opinion simulation tasks.
Harnessing the internal states of LLMs, SIREN outperforms traditional guard models while using a fraction of the parameters, revolutionizing harmful content detection.
Rollout design in LLM reinforcement learning is more than just sampling trajectories – it's a modular pipeline you can optimize for reliability, coverage, and cost.
Jointly training LLMs to reason and refine their answers unlocks significant performance gains, outperforming standard policy optimization by up to 11.5 points on AIME.