Search papers, labs, and topics across Lattice.
This paper formulates ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem, addressing the limitations of traditional RL approaches that optimize for a single ESG provider. By integrating a Preference Elicitation framework with Gaussian Processes, the authors enable practitioners to intuitively derive their utility functions through pairwise comparisons of portfolios based on Sharpe ratios and ESG scores. Empirical evaluations reveal significant regional variations in preference weights, demonstrating that European personas prioritize ESG alignment while Texas personas favor risk-adjusted returns, highlighting the adaptability of the proposed framework to real-world decision-making contexts.
Regional differences in ESG preferences can drastically alter portfolio optimization outcomes, revealing a critical need for tailored approaches in financial decision-making.
Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.