Search papers, labs, and topics across Lattice.
University of Queensland
2
0
3
A stealthy skill injection method that achieves an 89.3% success rate while evading detection in LLM agents reveals critical vulnerabilities in current safety mechanisms.
Chain-of-thought prompting makes large language models smarter, but it also makes them less safe, a problem this paper tackles by forcing models to think about safety *before* reasoning.