Search papers, labs, and topics across Lattice.
3
1
3
1
Safety intentions in AI agents can degrade over time, leading to unsafe actions and operational failures, but a new architectural layer can intercept these issues effectively.
Stop relying on aggregate safety scores: this new A-R behavioral space reveals *how* tool-augmented LLMs redistribute actions and refusals across risk contexts and autonomy levels.
AI students simultaneously overestimate explicit risks and underestimate scenario-based risks, leading to potentially reckless adoption despite stated awareness.