Search papers, labs, and topics across Lattice.
3
0
4
2
ColluSkill reveals that seemingly harmless skill combinations can orchestrate devastating attacks, achieving a staggering 96% success rate against current defenses.
A single poisoned skill can stealthily compromise LLM agents, activating only under specific conditions, raising alarms about the security of AI skill supply chains.
Forget broad brushstrokes: pinpointing and tweaking just 1% of an LLM's weights can slash attack success rates by over 50% or preserve safety during fine-tuning.