Search papers, labs, and topics across Lattice.
This paper introduces the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method, which utilizes a two-stage distillation learning framework to enhance control of remotely operated vehicles (ROVs) in the presence of ocean currents. By innovating in reward-function design and employing a privileged multi-encoder architecture, TSRCA-PPO achieves significant improvements in steady-state tracking, transient response, and energy efficiency compared to traditional cascaded P-PID controllers. Simulation results reveal that TSRCA-PPO reduces key performance metrics, including steady-state position error and settling time, to as low as 15.9% of the conventional method's values, showcasing its effectiveness in real-world applications.
TSRCA-PPO slashes control errors and energy consumption for ROVs, outperforming traditional methods by up to 93.5% in key metrics.
With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicles control. Nevertheless, existing end-to-end reinforcement-learningbased methods still face challenges in achieving optimal control under oceancurrent disturbances. In particular, there remains a lack of a unified control framework that can simultaneously achieve low steady-state tracking error, rapid transient response, energy-efficient operation, and smooth controlforce outputs under disturbances. To address the issue, this paper proposes the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method which learns a near-optimal strategy by a twostage distillation learning framework. The core innovations of this work lie in the reward-function design and the privileged multi-encoder architecture. Ablation studies validate the effectiveness of each module. Simulation results demonstrate that the proposed TSRCA-PPO method consistently outperforms the conventional cascaded P-PID controller across all evaluation metrics. Specifically, TSRCA-PPO reduces the steady-state position error, steady-state attitude error, settling time, energy index, and thrustsmoothness index to 42.7%, 76.5%, 10.6%, 93.5%, and 15.9% of the corresponding P-PID values, respectively.