CrackRobotics

Learning Algorithms, step 2

Q-learning

Learn action values from trial moves alone, exploring ε-greedily, then drive the arm over the wall with the greedy policy.

Temporal-difference learningExploration vs exploitationOff-policy learning

This step is part of CrackRobotics Pro

Pick & Place and the first step of every other lesson are free. Pro unlocks every other step for $39 a month, and you can cancel anytime.