Learning Algorithms
Let the robot work out what to do. Plan over the wall with value iteration, learn the same thing by trial and error with Q-learning, copy an expert with behaviour cloning, and teach the arm to throw a ball into the cup with REINFORCE.
Start the lesson0/4 steps done
What you'll learn
- Value iteration on a known MDP
- Tabular Q-learning with ε-greedy exploration
- Behaviour cloning with least squares
- Policy gradients with REINFORCE and a baseline
Before you start
You should be comfortable with basic Python: variables, loops and functions. We'll introduce the robotics and the maths as you go. The first visit downloads the physics engine and Python, about 20 MB, and later visits load from your browser's cache.
Steps
- 1Value iterationBack up the Bellman equation until the values stop changing, then follow the greedy policy to lift the arm over the wall.Markov decision processesBellman equationGreedy policiesOpen
- 2Q-learningLearn action values from trial moves alone, exploring ε-greedily, then drive the arm over the wall with the greedy policy.Temporal-difference learningExploration vs exploitationOff-policy learningPro
- 3Behaviour cloningCopy an expert: fit a policy to its demonstrations by least squares, then send it to targets it has never seen.Imitation learningLeast squaresFeaturesDistribution shiftPro
- 4Policy gradient (REINFORCE)Learn to throw a ball into the cup by trial and error: sample throws around your best guess and move towards the ones that land closer.Likelihood-ratio gradientBaselineLearning-rate decayPro