Search papers, labs, and topics across Lattice.
4
0
5
4
Evolving adversarial strategies can achieve nearly 100% success against state-of-the-art LLMs while maintaining robust defenses that adapt in real-time.
Frontier LLMs can be induced to generate biologically hazardous sequences, with attack success rates reaching up to 100%.
DataShield reveals that aligning consensus subspaces across multiple LLMs can drastically enhance safety by filtering out risky fine-tuning data more effectively than previous methods.
AgentDoG 1.5 proves you can achieve GPT-5.4-level agent safety with open-source models trained on just 1k samples, slashing deployment overhead by two orders of magnitude.