Search papers, labs, and topics across Lattice.
This paper presents a reinforcement learning approach for training sensorimotor policies that enable quadrotors to perform precise aggressive maneuvers, specifically traversing narrow gaps with significant tilt. The key innovation is an initialization strategy using model-based planning trajectories to guide RL exploration, addressing the challenge of sparse rewards in complex maneuver spaces. The resulting policies demonstrate successful sim-to-real transfer, enabling a quadrotor to navigate rectangular gaps with only 5 cm clearance and up to 90-degree tilt, even reacting to moving gaps without explicit training.
Drones can now thread the needle through narrow, tilted gaps with only onboard sensors, thanks to a new RL approach that leaps past exploration challenges.
Precise aggressive maneuvers with lightweight onboard sensors remains a key bottleneck in fully exploiting the maneuverability of drones. Such maneuvers are critical for expanding the systems'accessible area by navigating through narrow openings in the environment. Among the most relevant problems, a representative one is aggressive traversal through narrow gaps with quadrotors under SE(3) constraints, which require the quadrotors to leverage a momentary tilted attitude and the asymmetry of the airframe to navigate through gaps. In this paper, we achieve such maneuvers by developing sensorimotor policies directly mapping onboard vision and proprioception into low-level control commands. The policies are trained using reinforcement learning (RL) with end-to-end policy distillation in simulation. We mitigate the fundamental hardness of model-free RL's exploration on the restricted solution space with an initialization strategy leveraging trajectories generated by a model-based planner. Careful sim-to-real design allows the policy to control a quadrotor through narrow gaps with low clearances and high repeatability. For instance, the proposed method enables a quadrotor to navigate a rectangular gap at a 5 cm clearance, tilted at up to 90-degree orientation, without knowledge of the gap's position or orientation. Without training on dynamic gaps, the policy can reactively servo the quadrotor to traverse through a moving gap. The proposed method is also validated by training and deploying policies on challenging tracks of narrow gaps placed closely. The flexibility of the policy learning method is demonstrated by developing policies for geometrically diverse gaps, without relying on manually defined traversal poses and visual features.