Search papers, labs, and topics across Lattice.
2
6
4
5
Achieving top-tier performance in complex professional tasks with a model significantly smaller than its competitors reveals a new frontier in agentic intelligence.
RLHF's two-stage approach can statistically outperform DPO when learning from implicitly sparse rewards, challenging the narrative that end-to-end preference optimization is always superior.