Search papers, labs, and topics across Lattice.
This paper critiques existing multi-objective reinforcement learning (RL) approaches, specifically SER and ESR, for their failure to account for the interplay of non-linear utility effects occurring over varying timescales within the same decision problem. By providing both intuitive and numerical examples, the authors highlight a significant gap in the literature regarding how these different timescale effects impact user utility. The findings suggest that a more nuanced understanding of timescales in multi-objective planning could lead to improved decision-making strategies in RL applications.
Multi-objective RL methods overlook critical interactions between non-linear utility effects across different timescales, leading to suboptimal decision-making.
Time is of the essence when dealing with multiple reward signals and non-linear utility. In this paper we argue that the current main approaches in multi-objectiveRL (SER and ESR), and successor features, are insufficient. While each approach deals with non-linear effects on user utility on different timescales, none of them take into account that different effects happening on different timescales can happen within the same decision problem. We motivate that this can indeed be the case by an example, both intuitively and numerically, leading to a new perspective, and a significant and non-trivial gap in the literature.