Search papers, labs, and topics across Lattice.
3
0
5
0
Timing the entry of preference dimensions can lead to substantial performance gains in multi-preference alignment for LLMs.
Models trained with ACA-RL not only outperform on missing-premise tasks but also redefine how we evaluate reasoning under uncertainty in NLP.
Transforming agent failures into actionable recovery strategies, DARC enhances performance without bloating context, proving that less can be more in self-correction.