Search papers, labs, and topics across Lattice.
Affiliation:
3
0
3
Attack strategies can now evolve and adapt in real-time, leading to a staggering 48.6-point improvement against state-of-the-art models.
Early output distributions can be harnessed to drastically improve LLM safety assessments, cutting calibration errors by nearly 80%.
Long-tailed adversarial training can be revolutionized by leveraging confusion geometry to boost the robustness of vulnerable classes and sharpen critical decision boundaries.