Search papers, labs, and topics across Lattice.
3
0
2
This work introduces a reinforcement learning approach that forgoes predefined plans entirely, instead generating construction sequences adaptively as the structure is built, and evaluates the algorithm, HSAC, against the prior method hybrid-PPO (HPPO), demonstrating significantly higher asymptotic performance and good sample efficiency.
Safe sim-to-real transfer can be achieved with a new algorithm that guarantees near-optimal policies while minimizing real-world data collection risks.
Tight sample complexity guarantees for learning near-optimal policies in POMDPs are now possible, even from a single trajectory.