Skip to content

Reinforcement learning

Learning from rewards: Markov decision processes, tabular Q-learning worked by hand, deep Q-networks and learning to navigate.

This part builds on Foundations and Neural networks.

2 of 5 topics ready, listed in reading order
Markov decision processesStates, actions, rewards, returns, value functions and the Bellman equations.Planned
Q-learningTabular Q-learning on a small maze, worked by hand until it converges.ReadyDeep Q-networksReplay buffers, target networks and the CartPole benchmark.Ready
Game-playing agentsSelf-play on a small board game.Planned
Learning-based navigationA robot learning to reach goals from range scans, and the limits of sim-to-real.Planned