Search papers, labs, and topics across Lattice.
Affiliation:
3
1
4
0
Training methods reshape refusal mechanisms in language models, but no single approach achieves the trifecta of robustness, capability, and correctability.
Sequential preference optimization reveals a complex landscape where later training can enhance or degrade earlier preferences, depending on objective relationships.
LLM multi-agent systems can achieve significantly higher accuracy at a fraction of the cost by learning to selectively delegate tasks instead of relying on rigid orchestration.