Search papers, labs, and topics across Lattice.
4
0
7
2
Self-evolving rubric rewards can dramatically enhance audio reasoning in models, outperforming traditional methods by adapting to the model's evolving capabilities.
A stealthy skill injection method that achieves an 89.3% success rate while evading detection in LLM agents reveals critical vulnerabilities in current safety mechanisms.
Forget slow, expensive neural verifiers: this work shows a simple corpus lookup can provide faster, better rewards for RL fine-tuning of QA models.
LLMs can reason more robustly by fusing contextual hidden states with vocabulary embedding guidance, enabling dynamic switching between latent and explicit thinking.