Learning Algorithms, step 2
Q-learning
Learn action values from trial moves alone, exploring ε-greedily, then drive the arm over the wall with the greedy policy.
Temporal-difference learningExploration vs exploitationOff-policy learning
Learning Algorithms, step 2
Learn action values from trial moves alone, exploring ε-greedily, then drive the arm over the wall with the greedy policy.