Safe-Optimal Control for Motion Planning
A survey of safe-optimal control for motion planning based on reinforcement learning — covering optimal and safe control, game theory, sequential and demonstration-based learning, and motion planning. The sections below collect the key references for each topic.
Optimal Control
This section covers fundamental approaches to optimal control, including dynamic programming, linear programming, tree-based planning, control theory, and model predictive control.
Dynamic Programming
- (book) Dynamic Programming, Bellman R. (1957).
- (book) Dynamic Programming and Optimal Control, Volumes 1 and 2, Bertsekas D. (1995).
- (book) Markov Decision Processes - Discrete Stochastic Dynamic Programming, Puterman M. (1995).
- An Upper Bound on the Loss from Approximate Optimal-Value Functions, Singh S., Yee R. (1994).
- Stochastic optimization of sailing trajectories in an upwind regatta, Dalang R. et al. (2015).
Linear Programming
- (book) Markov Decision Processes - Discrete Stochastic Dynamic Programming, Puterman M. (1995).
REPSRelative Entropy Policy Search, Peters J. et al. (2010).
Tree-Based Planning
ExpectiMinimaxOptimal strategy in games with chance nodes, Melkó E., Nagy B. (2007).Sparse samplingA sparse sampling algorithm for near-optimal planning in large Markov decision processes, Kearns M. et al. (2002).MCTSEfficient Selectivity and Backup Operators in Monte-Carlo Tree Search, Rémi Coulom, SequeL (2006).UCTBandit based Monte-Carlo Planning, Kocsis L., Szepesvári C. (2006).- Bandit Algorithms for Tree Search, Coquelin P-A., Munos R. (2007).
OPDOptimistic Planning for Deterministic Systems, Hren J., Munos R. (2008).OLOPOpen Loop Optimistic Planning, Bubeck S., Munos R. (2010).SOOPOptimistic Planning for Continuous-Action Deterministic Systems, Buşoniu L. et al. (2011).OPSSOptimistic planning for sparsely stochastic systems, L. Buşoniu, R. Munos, B. De Schutter, and R. Babuska (2011).HOOTSample-Based Planning for Continuous ActionMarkov Decision Processes, Mansley C., Weinstein A., Littman M. (2011).HOLOPBandit-Based Planning and Learning inContinuous-Action Markov Decision Processes, Weinstein A., Littman M. (2012).BRUESimple Regret Optimization in Online Planning for Markov Decision Processes, Feldman Z. and Domshlak C. (2014).LGPLogic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning, Toussaint M. (2015). 🎞️AlphaGoMastering the game of Go with deep neural networks and tree search, Silver D. et al. (2016).AlphaGo ZeroMastering the game of Go without human knowledge, Silver D. et al. (2017).AlphaZeroMastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm, Silver D. et al. (2017).TrailBlazerBlazing the trails before beating the path: Sample-efficient Monte-Carlo planning, Grill J. B., Valko M., Munos R. (2017).MCTSnetsLearning to search with MCTSnets, Guez A. et al. (2018).ADISolving the Rubik’s Cube Without Human Knowledge, McAleer S. et al. (2018).OPC/SOPCContinuous-action planning for discounted infinite-horizon nonlinear optimal control with Lipschitz values, Buşoniu L., Pall E., Munos R. (2018).- Real-time tree search with pessimistic scenarios: Winning the NeurIPS 2018 Pommerman Competition, Osogami T., Takahashi T. (2019)
Control Theory
- (book) The Mathematical Theory of Optimal Processes, L. S. Pontryagin, Boltyanskii V. G., Gamkrelidze R. V., and Mishchenko E. F. (1962).
- (book) Constrained Control and Estimation, Goodwin G. (2005).
PI²A Generalized Path Integral Control Approach to Reinforcement Learning, Theodorou E. et al. (2010).PI²-CMAPath Integral Policy Improvement with Covariance Matrix Adaptation, Stulp F., Sigaud O. (2010).iLQGA generalized iterative LQG method for locally-optimal feedback control of constrained nonlinear stochastic systems, Todorov E. (2005). :octocat:iLQG+Synthesis and stabilization of complex behaviors through online trajectory optimization, Tassa Y. (2012).
Model Predictive Control
- (book) Model Predictive Control, Camacho E. (1995).
- (book) Predictive Control With Constraints, Maciejowski J. M. (2002).
- Linear Model Predictive Control for Lane Keeping and Obstacle Avoidance on Low Curvature Roads, Turri V. et al. (2013).
MPCCOptimization-based autonomous racing of 1:43 scale RC cars, Liniger A. et al. (2014). 🎞️ | 🎞️MIQPOptimal trajectory planning for autonomous driving integrating logical constraints: An MIQP perspective, Qian X., Altché F., Bender P., Stiller C. de La Fortelle A. (2016).
Safe Control
This section covers approaches to ensuring safety in control systems, including robust control, risk-averse control, value-constrained control, state-constrained control and stability, and uncertain dynamical systems.
Robust Control
- Minimax analysis of stochastic problems, Shapiro A., Kleywegt A. (2002).
Robust DPRobust Dynamic Programming, Iyengar G. (2005).- Robust Planning and Optimization, Laumanns M. (2011). (lecture notes)
- Robust Markov Decision Processes, Wiesemann W., Kuhn D., Rustem B. (2012).
- Safe and Robust Learning Control with Gaussian Processes, Berkenkamp F., Schoellig A. (2015). 🎞️
Tube-MPPIRobust Sampling Based Model Predictive Control with Sparse Objective Information, Williams G. et al. (2018). 🎞️- Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning, Lukas Bronke et al. (2021). :octocat:
Risk-Averse Control
- A Comprehensive Survey on Safe Reinforcement Learning, García J., Fernández F. (2015).
RA-QMDPRisk-averse Behavior Planning for Autonomous Driving under Uncertainty, Naghshvar M. et al. (2018).StoROOX-Armed Bandits: Optimizing Quantiles and Other Risks, Torossian L., Garivier A., Picheny V. (2019).- Worst Cases Policy Gradients, Tang Y. C. et al. (2019).
- Model-Free Risk-Sensitive Reinforcement Learning, Delétang G. et al. (2021).
- Optimal Thompson Sampling strategies for support-aware CVaR bandits, Baudry D., Gautron R., Kaufmann E., Maillard O. (2021).
Value-Constrained Control
ICSWill the Driver Seat Ever Be Empty?, Fraichard T. (2014).SafeOPTSafe Controller Optimization for Quadrotors with Gaussian Processes, Berkenkamp F., Schoellig A., Krause A. (2015). 🎞️ :octocat:SafeMDPSafe Exploration in Finite Markov Decision Processes with Gaussian Processes, Turchetta M., Berkenkamp F., Krause A. (2016). :octocat:RSSOn a Formal Model of Safe and Scalable Self-driving Cars, Shalev-Shwartz S. et al. (2017).CPOConstrained Policy Optimization, Achiam J., Held D., Tamar A., Abbeel P. (2017). :octocat:RCPOReward Constrained Policy Optimization, Tessler C., Mankowitz D., Mannor S. (2018).BFTQA Fitted-Q Algorithm for Budgeted MDPs, Carrara N. et al. (2018).SafeMPCLearning-based Model Predictive Control for Safe Exploration, Koller T, Berkenkamp F., Turchetta M. Krause A. (2018).CCEConstrained Cross-Entropy Method for Safe Reinforcement Learning, Wen M., Topcu U. (2018). :octocat:LTL-RLReinforcement Learning with Probabilistic Guarantees for Autonomous Driving, Bouton M. et al. (2019).- Safe Reinforcement Learning with Scene Decomposition for Navigating Complex Urban Environments, Bouton M. et al. (2019). :octocat:
- Batch Policy Learning under Constraints, Le H., Voloshin C., Yue Y. (2019).
- Value constrained model-free continuous control, Bohez S. et al (2019). 🎞️
- Safely Learning to Control the Constrained Linear Quadratic Regulator, Dean S. et al (2019).
- Learning to Walk in the Real World with Minimal Human Effort, Ha S. et al. (2020) 🎞️
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods, Stooke A., Achiam J., Abbeel P. (2020). :octocat:
Envelope MOQ-LearningA Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation, Yang R. et al (2019).
State-Constrained Control and Stability
HJI-reachabilitySafe learning for control: Combining disturbance estimation, reachability analysis and reinforcement learning with systematic exploration, Heidenreich C. (2017).MPC-HJIOn Infusing Reachability-Based Safety Assurance within Probabilistic Planning Frameworks for Human-Robot Vehicle Interactions, Leung K. et al. (2018).- A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems, Fisac J. et al (2017). 🎞️
- Safe Model-based Reinforcement Learning with Stability Guarantees, Berkenkamp F. et al. (2017).
Lyapunov-NetSafe Interactive Model-Based Learning, Gallieri M. et al. (2019).- Enforcing robust control guarantees within neural network policies, Donti P. et al. (2021). :octocat:
ATACOMRobot Reinforcement Learning on the Constraint Manifold, Liu P. et al (2021).
Uncertain Dynamical Systems
- Simulation of Controlled Uncertain Nonlinear Systems, Tibken B., Hofer E. (1995).
- Trajectory computation of dynamic uncertain systems, Adrot O., Flaus J-M. (2002).
- Simulation of Uncertain Dynamic Systems Described By Interval Models: a Survey, Puig V. et al. (2005).
- Design of interval observers for uncertain dynamical systems, Efimov D., Raïssi T. (2016).
Game Theory
This section covers game-theoretic approaches to multi-agent control and decision-making.
- Hierarchical Game-Theoretic Planning for Autonomous Vehicles, Fisac J. et al. (2018).
- Efficient Iterative Linear-Quadratic Approximations for Nonlinear Multi-Player General-Sum Differential Games, Fridovich-Keil D. et al. (2019). 🎞️
Sequential Learning
This section covers sequential learning approaches including multi-armed bandits, contextual bandits, best arm identification, and black-box optimization.
- Prediction, Learning and Games, Cesa-Bianchi N., Lugosi G. (2006).
Multi-Armed Bandit
TSOn the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples, Thompson W. (1933).- Exploration and Exploitation in Organizational Learning, March J. (1991).
UCB1 / UCB2Finite-time Analysis of the Multiarmed Bandit Problem, Auer P., Cesa-Bianchi N., Fischer P. (2002).Empirical Bernstein / UCB-VExploration-exploitation tradeoff using variance estimates in multi-armed bandits, Audibert J-Y, Munos R., Szepesvari C. (2009).- Empirical Bernstein Bounds and Sample Variance Penalization, Maurer A., Ponti M. (2009).
- An Empirical Evaluation of Thompson Sampling, Chapelle O., Li L. (2011).
kl-UCBThe KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond, Garivier A., Cappé O. (2011).KL-UCBKullback-Leibler Upper Confidence Bounds for Optimal Sequential Allocation, Cappé O. et al. (2013).IDSInformation Directed Sampling and Bandits with Heteroscedastic Noise Kirschner J., Krause A. (2018).
Contextual Bandits
LinUCBA Contextual-Bandit Approach to Personalized News Article Recommendation, Li L. et al. (2010).OFULImproved Algorithms for Linear Stochastic Bandits, Abbasi-yadkori Y., Pal D., Szepesvári C. (2011).- Contextual Bandits with Linear Payoff Functions, Chu W. et al. (2011).
- Self-normalization techniques for streaming confident regression, Maillard O.-A. (2017).
- Learning from Delayed Outcomes via Proxies with Applications to Recommender Systems Mann T. et al. (2018). (prediction setting)
- Weighted Linear Bandits for Non-Stationary Environments, Russac Y. et al. (2019).
- Linear bandits with Stochastic Delayed Feedback, Vernade C. et al. (2020).
Best Arm Identification
Successive EliminationAction Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems, Even-Dar E. et al. (2006).LUCBPAC Subset Selection in Stochastic Multi-armed Bandits, Kalyanakrishnan S. et al. (2012).UGapEBest Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence, Gabillon V., Ghavamzadeh M., Lazaric A. (2012).Sequential HalvingAlmost Optimal Exploration in Multi-Armed Bandits, Karnin Z. et al (2013).M-LUCB / M-RacingMaximin Action Identification: A New Bandit Framework for Games, Garivier A., Kaufmann E., Koolen W. (2016).Track-and-StopOptimal Best Arm Identification with Fixed Confidence, Garivier A., Kaufmann E. (2016).LUCB-microStructured Best Arm Identification with Fixed Confidence, Huang R. et al. (2017).
Black-box Optimization
GP-UCBGaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design, Srinivas N., Krause A., Kakade S., Seeger M. (2009).HOOX–Armed Bandits, Bubeck S., Munos R., Stoltz G., Szepesvari C. (2009).DOO/SOOOptimistic Optimization of a Deterministic Function without the Knowledge of its Smoothness, Munos R. (2011).StoOOFrom Bandits to Monte-Carlo Tree Search: The Optimistic Principle Applied to Optimization and Planning, Munos R. (2014).StoSOOStochastic Simultaneous Optimistic Optimization, Valko M., Carpentier A., Munos R. (2013).POOBlack-box optimization of noisy functions with unknown smoothness, Grill J-B., Valko M., Munos R. (2015).EI-GPBayesian Optimization in AlphaGo, Chen Y. et al. (2018)
Reinforcement Learning
This section covers the comprehensive landscape of reinforcement learning, from theoretical foundations to practical applications in control and decision-making.
- Reinforcement learning: A survey, Kaelbling L. et al. (1996).
Theory
Generative Model
QVIOn the Sample Complexity of Reinforcement Learning with a Generative Model, Azar M., Munos R., Kappen B. (2012).- Model-Based Reinforcement Learning with a Generative Model is Minimax Optimal, Agarwal A. et al. (2019).
Policy Gradient
- Policy Gradient Methods for Reinforcement Learning with Function Approximation, Sutton R. et al (2000).
- Approximately Optimal Approximate Reinforcement Learning, Kakade S., Langford J. (2002).
- On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift, Agarwal A. et al. (2019)
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning, Agarwal A. et al. (2020)
- Is the Policy Gradient a Gradient?, Nota C., Thomas P. S. (2020).
Linear Systems
- PAC Adaptive Control of Linear Systems, Fiechter C.-N. (1997)
OFU-LQRegret Bounds for the Adaptive Control of Linear Quadratic Systems, Abbasi-Yadkori Y., Szepesvari C. (2011).TS-LQImproved Regret Bounds for Thompson Sampling in Linear Quadratic Control Problems, Abeille M., Lazaric A. (2018).
Value-based Methods
DQNPlaying Atari with Deep Reinforcement Learning, Mnih V. et al. (2013). 🎞️DDQNDeep Reinforcement Learning with Double Q-learning, van Hasselt H., Silver D. et al. (2015).RainbowRainbow: Combining Improvements in Deep Reinforcement Learning, Hessel M. et al. (2017).
Policy-based Methods
Policy Gradient
REINFORCESimple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning, Williams R. (1992).TRPOTrust Region Policy Optimization, Schulman J. et al. (2015). 🎞️PPOProximal Policy Optimization Algorithms, Schulman J. et al. (2017). 🎞️
Actor-Critic
DDPGContinuous Control With Deep Reinforcement Learning, Lillicrap T. et al. (2015).A3CAsynchronous Methods for Deep Reinforcement Learning, Mnih V. et al 2016.SACSoft Actor-Critic : Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Haarnoja T. et al. (2018). 🎞️
Model-based Methods
PILCOPILCO: A Model-Based and Data-Efficient Approach to Policy Search, Deisenroth M., Rasmussen C. (2011).MPPIInformation Theoretic MPC for Model-Based Reinforcement Learning, Williams G. et al. (2017). :octocat: 🎞️MuZeroMastering Atari, Go, Chess and Shogi by Planning with a Learned Model, Schrittwiese J. et al. (2019). :octocat:
Exploration
HERHindsight Experience Replay, Andrychowicz M. et al. (2017). 🎞️RNDExploration by Random Network Distillation, Burda Y. et al. (OpenAI) (2018). 🎞️Go-ExploreGo-Explore: a New Approach for Hard-Exploration Problems, Ecoffet A. et al. (Uber) (2018). 🎞️
Multi-agent RL
MADDPGMulti-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, Lowe R. et al (2017). :octocat:FTWHuman-level performance in first-person multiplayer games with population-based deep reinforcement learning, Jaderberg M. et al. (2018). 🎞️MAPPOThe Surprising Effectiveness of MAPPO in Cooperative, Multi-Agent Games, Yu C. et al. (2021). :octocat:
Safe Reinforcement Learning
- A Comprehensive Survey on Safe Reinforcement Learning, García J., Fernández F. (2015).
CPOConstrained Policy Optimization, Achiam J., Held D., Tamar A., Abbeel P. (2017). :octocat:- Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning, Lukas Bronke et al. (2021). :octocat:
Transfer Learning and Meta-Learning
MAMLModel-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, Finn C., Abbeel P., Levine S. (2017). 🎞️- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots, Tan J. et al. (2018). 🎞️
- Learning Dexterous In-Hand Manipulation, OpenAI (2018). 🎞️
Hierarchical RL
OCThe Option-Critic Architecture, Bacon P-L., Harb J., Precup D. (2016).FuNsFeUdal Networks for Hierarchical Reinforcement Learning, Vezhnevets A. et al. (2017).DeepLocoDeepLoco: Dynamic Locomotion Skills Using Hierarchical Deep Reinforcement Learning, Peng X. et al. (2017). 🎞️
Offline RL
CQLConservative Q-Learning for Offline Reinforcement Learning, Kumar A. et al. (2020).- Decision Transformer: Reinforcement Learning via Sequence Modeling, Chen L., Lu K. et al. (2021). :octocat:
Note
This section provides a comprehensive overview of reinforcement learning approaches relevant to safe and optimal control. For the complete list of papers and more detailed subsections, please refer to the original survey document.
Learning from Demonstrations
This section covers approaches to learning control policies from expert demonstrations, including imitation learning and inverse reinforcement learning.
Imitation Learning
DAggerA Reduction of Imitation Learning and Structured Predictionto No-Regret Online Learning, Ross S., Gordon G., Bagnell J. A. (2011).GAILGenerative Adversarial Imitation Learning, Ho J., Ermon S. (2016).DQfDLearning from Demonstrations for Real World Reinforcement Learning, Hester T. et al. (2017). 🎞️BranchedEnd-to-end Driving via Conditional Imitation Learning, Codevilla F. et al. (2017). 🎞️DeepMimicDeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills, Peng X. B. et al. (2018). 🎞️
Applications to Autonomous Driving
- ALVINN, an autonomous land vehicle in a neural network, Pomerleau D. (1989).
- End to End Learning for Self-Driving Cars, Bojarski M. et al. (2016). 🎞️
- Imitating Driver Behavior with Generative Adversarial Networks, Kuefler A. et al. (2017).
PS-GAILMulti-Agent Imitation Learning for Driving Simulation, Bhattacharyya R. et al. (2018). 🎞️ :octocat:
Inverse Reinforcement Learning
ProjectionApprenticeship learning via inverse reinforcement learning, Abbeel P., Ng A. (2004).MEIRLMaximum Entropy Inverse Reinforcement Learning, Ziebart B. et al. (2008).CIOCContinuous Inverse Optimal Control with Locally Optimal Examples, Levine S., Koltun V. (2012). 🎞️GCLGuided Cost Learning: Deep Inverse Optimal Control via Policy Optimization, Finn C. et al. (2016). 🎞️
Applications to Autonomous Driving
- Apprenticeship Learning for Motion Planning, with Application to Parking Lot Navigation, Abbeel P. et al. (2008).
- Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior, Ziebart B. et al. (2008).
- Learning Autonomous Driving Styles and Maneuvers from Expert Demonstration, Silver D. et al. (2012).
- Watch This: Scalable Cost-Function Learning for Path Planning in Urban Environments, Wulfmeier M. (2016). 🎞️
Motion Planning
This section covers fundamental approaches to motion planning, including search-based methods, sampling-based methods, optimization approaches, reactive methods, and their applications.
Search
DijkstraA Note on Two Problems in Connexion with Graphs, Dijkstra E. W. (1959).A*A Formal Basis for the Heuristic Determination of Minimum Cost Paths, Hart P. et al. (1968).- Planning Long Dynamically-Feasible Maneuvers For Autonomous Vehicles, Likhachev M., Ferguson D. (2008).
- Optimal Trajectory Generation for Dynamic Street Scenarios in a Frenet Frame, Werling M., Kammel S. (2010). 🎞️
- Motion Planning under Uncertainty for On-Road Autonomous Driving, Xu W. et al. (2014).
- Monte Carlo Tree Search for Simulated Car Racing, Fischer J. et al. (2015). 🎞️
Sampling
RRT*Sampling-based Algorithms for Optimal Motion Planning, Karaman S., Frazzoli E. (2011). 🎞️LQG-MPLQG-MP: Optimized Path Planning for Robots with Motion Uncertainty and Imperfect State Information, van den Berg J. et al. (2010).- Motion Planning under Uncertainty using Differential Dynamic Programming in Belief Space, van den Berg J. et al. (2011).
- Rapidly-exploring Random Belief Trees for Motion Planning Under Uncertainty, Bry A., Roy N. (2011).
PRM-RLPRM-RL: Long-range Robotic Navigation Tasks by Combining Reinforcement Learning and Sampling-based Planning, Faust A. et al. (2017).
Optimization
- Trajectory planning for Bertha - A local, continuous method, Ziegler J. et al. (2014).
- Learning Attractor Landscapes for Learning Motor Primitives, Ijspeert A. et al. (2002).
- Online Motion Planning based on Nonlinear Model Predictive Control with Non-Euclidean Rotation Groups, Rösmann C. et al (2020). :octocat:
Reactive
PFReal-time obstacle avoidance for manipulators and mobile robots, Khatib O. (1986).VFHThe Vector Field Histogram - Fast Obstacle Avoidance For Mobile Robots, Borenstein J. (1991).VFH+VFH+: Reliable Obstacle Avoidance for Fast Mobile Robots, Ulrich I., Borenstein J. (1998).Velocity ObstaclesMotion planning in dynamic environments using velocity obstacles, Fiorini P., Shillert Z. (1998).
Architecture and Applications
- A Review of Motion Planning Techniques for Automated Vehicles, González D. et al. (2016).
- A Survey of Motion Planning and Control Techniques for Self-driving Urban Vehicles, Paden B. et al. (2016).
- Autonomous driving in urban environments: Boss and the Urban Challenge, Urmson C. et al. (2008).
- The MIT-Cornell collision and why it happened, Fletcher L. et al. (2008).
- Making bertha drive-an autonomous journey on a historic route, Ziegler J. et al. (2014).