Cross-Observability Optimistic-Pessimistic Safe Reinforcement Learning for Interactive Motion Planning With Visual Occlusion

This study focuses on the motion planning and risk evaluation of unprotected left turns at occluded intersections for autonomous vehicles. In this paper, we present an interactive motion planning controller that combines Cross-Observability Optimistic-Pessimistic Safe Reinforcement Learning (COOP-SR...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on intelligent transportation systems 2024-11, Vol.25 (11), p.17602-17613
Hauptverfasser:	Hou, Xiaohui, Gan, Minggang, Wu, Wei, Ji, Yuan, Zhao, Shiyue, Chen, Jie
Format:	Artikel
Sprache:	eng
Schlagworte:	Autonomous vehicles Motion planning Planning Reinforcement learning risk evaluation Safety Uncertainty Vehicle dynamics visual occlusion Visualization
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	This study focuses on the motion planning and risk evaluation of unprotected left turns at occluded intersections for autonomous vehicles. In this paper, we present an interactive motion planning controller that combines Cross-Observability Optimistic-Pessimistic Safe Reinforcement Learning (COOP-SRL) and Nonlinear Model Predictive Control (NMPC), with consideration of the uncertain potential risk of occluded zone, the trade-off between safety and efficiency, and the dynamic interaction between vehicles. The proposed COOP-SRL algorithm integrates fully and partially observable policies through cross-observability soft imitation learning to leverage the expert guidance and improve learning efficiency. Moreover, the optimistic exploration policy and pessimism safe constraint are adopted to provide an adaptive safe strategy without hindering the exploration during learning process. Finally, the evaluations of the proposed controller were conducted in occluded intersection scenarios with various traffic density level, which indicate that the proposed method outperforms both the optimization-based and learning-based baselines in qualitative and quantitative indexes.
ISSN:	1524-9050 1558-0016
DOI:	10.1109/TITS.2024.3443397