Search papers, labs, and topics across Lattice.
2
0
5
Modern LLMs rarely refuse to discuss restricted content, instead opting for nuanced warnings that reveal a significant shift in content moderation strategies.
LLMs' chain-of-thought explanations often fail to reflect the true drivers of their decisions, and this benchmark reveals that closed-source models are particularly opaque, with monitorability dropping by up to 30% under stress.