Search papers, labs, and topics across Lattice.
Vrije Universiteit Brussel
2
0
4
Summaries may seem helpful, but they often mislead users about correctness compared to the full reasoning trace, especially when prompts are withheld.
Perturbation techniques can exploit subtle internal representations in LLMs, revealing vulnerabilities that could compromise model safety.