Reinforcement learning¶
Learning from rewards: Markov decision processes, tabular Q-learning worked by hand, deep Q-networks and learning to navigate.
This part builds on Foundations and Neural networks.
2 of 5 topics ready, listed in reading order
Markov decision processesStates, actions, rewards, returns, value functions and the Bellman equations.Planned
Q-learningTabular Q-learning on a small maze, worked by hand until it converges.ReadyDeep Q-networksReplay buffers, target networks and the CartPole benchmark.ReadyGame-playing agentsSelf-play on a small board game.Planned
Learning-based navigationA robot learning to reach goals from range scans, and the limits of sim-to-real.Planned