Search papers, labs, and topics across Lattice.
Affiliation:
8
0
11
Standard single-turn RL produces socially myopic agents, but dynamically scheduling rewards from early relationship-building to mid-dialogue task execution drives a 9.2 percentage point gain in multi-turn goal achievement.
Safety compliance can vary by over 28 points among LLMs with similar predictive accuracy, revealing hidden risks in flight predictions.
DECAF achieves near-optimal unlearning performance while maintaining efficiency, effectively neutralizing clustering attacks that threaten data privacy.
HealthClaw boosts answer accuracy for personal health management by over 45% while enhancing privacy protection in AI interactions.
Adversarial documents can not only mislead deep research agents but also shift poisoned content from overt framing into seemingly factual premises, complicating detection.
Adaptive weighting in model merging can drastically improve multilingual reasoning performance, outperforming traditional methods across 21 languages.
MLLMs exhibit alarming Stochastic Collapse, failing to maintain randomness even under explicit random instructions, which could undermine their utility in diverse applications.
Frontier models can't build playable games in one shot, but a closed-loop system using GUI agents to playtest and refine code achieves a 66.8% success rate, proving that game generation needs to be a conversation, not a translation.