Course
Deep Reinforcement Learning
A comprehensive course on modern reinforcement learning — from bandits and MDPs to DQN, policy gradients, and model-based methods.
A comprehensive course on modern reinforcement learning.
This series builds Deep RL from the ground up, in four parts:
- Tabular RL — bandits, Markov decision processes and dynamic programming, Monte Carlo control, temporal-difference learning, function approximation, and the deep-network building blocks.
- Model-free RL — DQN and its variants, policy gradient, actor-critic (A3C), DDPG, PPO, ACER, and maximum-entropy RL (SAC).
- Model-based RL — model-based RL, planning with learned world models (PlaNet, Dreamer), importance sampling, MPC, AlphaGo, and natural gradients.
- Outlook — current research directions: inverse RL, intrinsic motivation, hierarchical, meta, offline, and multi-agent RL.
Read the lectures in order below. Where a slide deck exists, it’s linked at the top of the lecture.
Recommended readings
- Sutton & Barto (2018), Reinforcement Learning: An Introduction — http://incompleteideas.net/book/the-book-2nd.html
- Sergey Levine, CS294 Deep RL (UC Berkeley) — http://rll.berkeley.edu/deeprlcourse/
- David Silver, Reinforcement Learning (UCL) — http://www0.cs.ucl.ac.uk/staff/d.silver/web/Teaching.html
- 01 Introduction 10 min
- 02 Sampling and Bandits 15 min
- 03 Markov Decision Process 23 min
- 04 Monte Carlo methods 10 min
- 05 Temporal Difference learning 17 min
- 06 Function approximation 7 min
- 07 Deep learning 12 min
- 08 Deep Q-network (DQN) 10 min
- 09 DQN variants (Rainbow) 21 min
- 10 Policy Gradient methods 13 min
- 11 Advantage Actor-Critic (A3C) 14 min
- 12 Deep Deterministic Policy Gradient (DDPG) 15 min
- 13 Policy optimization (TRPO, PPO) 14 min
- 14 Actor-Critic with Experience Replay (ACER) 10 min
- 15 Maximum Entropy RL (SAC) 12 min
- 16 Model-based RL 4 min
- 17 Model-based-augmented model-free RL (Dyna-Q, I2A) 7 min
- 18 Off-policy Actor-Critic 20 min
- 19 Planning (MPC, TDM) 9 min
- 20 World models, Dreamer 8 min
- 21 AlphaGo 11 min
- 22 Natural gradients 12 min
- 23 Intrinsic motivation 1 min