Search papers, labs, and topics across Lattice.
This paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that quantifies an agent's caution based on epistemic uncertainty by utilizing a Bayesian posterior over unknown dynamics and a Wasserstein ambiguity set. The approach allows for a continuous interpolation between worst-case robustness and risk-neutral total-reward maximization as evidence accumulates, ensuring that the agent's behavior adapts to its learning progress. The authors establish a "Safety Sandwich" theorem, demonstrating that RATTL's value consistently lies between uninformed robust values and optimal values, converging as the agent's knowledge improves.
RATTL allows agents to dynamically balance caution and reward maximization, adapting their decision-making as they learn about their environment.
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.