Search papers, labs, and topics across Lattice.
3
0
5
2
Evolving adversarial strategies can achieve nearly 100% success against state-of-the-art LLMs while maintaining robust defenses that adapt in real-time.
DataShield reveals that aligning consensus subspaces across multiple LLMs can drastically enhance safety by filtering out risky fine-tuning data more effectively than previous methods.
Forget broad brushstrokes: pinpointing and tweaking just 1% of an LLM's weights can slash attack success rates by over 50% or preserve safety during fine-tuning.