Search papers, labs, and topics across Lattice.
This paper introduces a noise-robust elicit-to-optimize framework that combines inverse reinforcement learning (IRL) and reinforcement learning (RL) to effectively elicit agents' risk preferences and optimize policies under distortion riskmetrics. The authors present an adaptive Bayesian IRL method that accurately infers latent risk objectives from noisy decision data, proving convergence rates that enhance the robustness of elicitation. Additionally, they develop a model-free RL algorithm that leverages quantile neural networks to optimize policies across a range of risk objectives, demonstrating superior performance in complex financial scenarios.
Eliciting risk preferences from noisy decisions can now be done with unprecedented accuracy, enabling more effective policy optimization in uncertain environments.
We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents'risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents'latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.