Search papers, labs, and topics across Lattice.
This paper addresses the challenge of stochastic control in multivariate Hawkes-driven stochastic differential equations, which are inherently non-Markovian due to their path-dependent memory. The authors introduce a finite-dimensional Markovianization procedure that approximates these processes using mixtures of exponential kernels, proving convergence to the original non-Markovian dynamics. They then develop a continuous-time deterministic policy gradient learning algorithm, Hawkes-CT DDPG, which effectively optimizes non-Markovian Hawkes processes using only event times and selected decay filters, outperforming traditional discrete-time methods across various kernel types.
Continuous-time reinforcement learning can effectively tackle non-Markovian Hawkes processes, outperforming discrete methods in optimization tasks.
We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation of the problem, called Hawkes-CT DDPG. We propose a model-free algorithm to solve the non-Markovian Hawkes-driven optimization by observing only the event times of the process, the realization of the solution to the SDE, and a chosen set of decay filters, while the Hawkes kernel coefficients remain unknown. We compare our continuous time reinforcement learning Hawkes-CT DDPG method with discrete time reinforcement learning techniques under three different types of kernels: simple exponential, Erlang, and power-law kernels.