Search papers, labs, and topics across Lattice.
6
0
11
13
BabelSteering boosts harmful request refusals across multiple languages by an average of 11 percentage points, all while maintaining task performance.
Faithfulness and safety in LRMs are at odds, with one model achieving high accuracy but failing to reject unsafe reasoning, while another sacrifices accuracy for improved safety.
Frontier LLMs excel in social deduction games, but most fail to sustain deception, with retention rates plummeting below 50%.
Misinformation can persist through multi-agent interactions, but structured debate can significantly reduce its impact on decision-making performance.
LLMs alone can't capture the nuances of mathematical research, but injecting aspect-aware information into a heterogeneous GNN unlocks surprisingly effective paper recommendations.
TCP and QUIC can achieve NAT traversal success rates statistically indistinguishable from UDP, overturning long-held beliefs about P2P networking.