Search papers, labs, and topics across Lattice.
10
0
14
0
AdvSafe enables LRMs to achieve jailbreak robustness with minimal utility loss by teaching them to understand and counteract adversarial threats intrinsically.
RoMeRL achieves an 80% reduction in the Cold-Q ratio while enhancing feedback density sixfold, revolutionizing how LLM agents manage memory and rewards.
AlphaSchema reveals that systematic exploration of trading semantics can significantly enhance alpha mining effectiveness, yielding robust predictive performance across various LLMs.
Thinking enhances decision-making in LLMs by improving action based on current evidence, but doesn鈥檛 drive them to seek more information.
Role-playing agents can achieve human-like character fidelity by integrating psychology-grounded reasoning with advanced reinforcement learning techniques.
GraphPO slashes redundancy in reasoning model training, enabling more efficient exploration and improved performance on complex tasks.
Forget tedious, brittle automation scripts: RL-powered GUI agents are showing signs of "System 2" reasoning without explicit supervision, hinting at a future of truly intelligent digital inhabitants.
Highlighting pivotal evidence can boost LLM performance without altering the original context, leading to substantial improvements in reasoning tasks.
Serverless computing slashes RLHF costs by up to 45% while boosting speed by 35%, thanks to a new framework that dynamically scales resources and avoids redundant computation.
Personalized federated learning can boost VLN performance by up to 7.8% in trajectory fidelity and converge 1.38x faster, by selectively fusing parameters in environment-sensitive layers.