Problems · Problem 21 · Learning · Easy
Behaviour cloning
Copy an expert: fit a policy to its demonstrations by least squares, then send it to targets it has never seen.
Builds on Least squares: fitting noisy data, from the free Foundations.
Write and run this problem in the simulator with ProWhat it computes
Behaviour cloning learns a policy by copying an expert: record what it saw, , and what it did, , then fit . It is plain supervised learning. The expert's actions are the labels, so no reward and no trial and error are needed.
Least squares on features
Make the policy linear in features: . Stack the rows into and the actions into , and pick the best fit:
X = [features(o) for o in obs] # N x k
W = least-squares solution of X W = actions # k x 4
policy(o) = features(o) @ W
Choosing the features
An observation is : the joint angles, then = target − gripper (m). Use : then the policy is exactly zero on the target, where the expert stops too. Regress on all seven numbers and the terms don't vanish at , so the arm parks about 1 cm short (1–7 of 20 targets instead of 20). Try it and watch the plot.
Why it can fail
The policy only learned the states the expert visited. Its small mistakes lead it to states nobody demonstrated, where it errs more: errors compound (distribution shift). DAgger fixes this: run the policy, have the expert label the states it reaches, add them to the data and refit.
Tools
rl.expert_demos(40, seed=0) returns (obs, actions), 2,400 rows of 7 and 4 numbers. np.linalg.lstsq(X, Y, rcond=None)[0] solves least squares.
See it in context
The expert is the damped least-squares step from the Kinematics set's Numerical inverse kinematics, run as a controller: .
Your task
Implement features(obs) for one observation and fit(obs, actions), which returns the policy: a function from one observation to 4 joint velocities (rad/s).
It must bring the gripper within 1 cm of at least 16 of 20 targets it has never seen. The grader also tests fit on made-up data.