Policy Search for Model Predictive Control With Application to Agile Drone Flight

Policy search and model predictive control (MPC) are two different paradigms for robot control: policy search has the strength of automatically learning complex policies using experienced data, and MPC can offer optimal control performance using models and trajectory optimization. An open research q...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on robotics 2022-08, Vol.38 (4), p.2114-2130
Hauptverfasser:	Song, Yunlong, Scaramuzza, Davide
Format:	Artikel
Sprache:	eng
Schlagworte:	Automation Controllers Drones Learning Learning agile flight Logic gates model predictive control (MPC) Neural networks Optimal control Policies Predictive control Predictive models Probabilistic logic reinforcement learning (RL) Robot control Robust control Searching Task analysis Trajectory optimization Vehicle dynamics
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Policy search and model predictive control (MPC) are two different paradigms for robot control: policy search has the strength of automatically learning complex policies using experienced data, and MPC can offer optimal control performance using models and trajectory optimization. An open research question is how to leverage and combine the advantages of both approaches. In this article, we provide an answer by using policy search for automatically choosing high-level decision variables for MPC, which leads to a novel policy-search-for-model-predictive-control framework . Specifically, we formulate the MPC as a parameterized controller, where the hard-to-optimize decision variables are represented as high-level policies. Such a formulation allows optimizing policies in a self-supervised fashion. We validate this framework by focusing on a challenging problem in agile drone flight: flying a quadrotor through fast-moving gates. Experiments show that our controller achieves robust and real-time control performance in both simulation and the real world. The proposed framework offers a new perspective for merging learning and control.
ISSN:	1552-3098 1941-0468
DOI:	10.1109/TRO.2022.3141602