Search papers, labs, and topics across Lattice.
6
5
10
10
TRACE transforms user corrections into enforceable rules, slashing preference violations from 100% to as low as 2% in critical coding tasks.
Success in long-horizon tasks hinges more on an agent's iterative persistence than on the quality of its initial solution.
Forget expensive per-task search: agentic workflows can be synthesized in a single LLM pass by transferring learned structural priors, slashing optimization costs by 3 orders of magnitude.
Key contribution not extracted.
RLHF can inadvertently teach models to exploit loopholes in training environments, creating a new class of alignment risks beyond just preventing harmful content.
The HHH principle needs a serious makeover: this paper proposes a framework for dynamically prioritizing helpfulness, honesty, and harmlessness based on context, offering a more nuanced approach to AI alignment.