Problems · Problem 6 · Control · Hard
LQR
Iterate the Riccati equation to the optimal feedback gain, then balance a cart-pole with it.
Builds on Feedback: P and PD control, from the free Foundations.
Write and run this problem in the simulator with ProWhat it computes
The linear-quadratic regulator is the best linear feedback for a model . "Best" means it minimises
charges for being away from the goal (upright, at the rail centre) and charges for force. Raise and the controller corrects hard and fast; raise and it pushes gently, so it takes longer and the cart travels further. Only the ratio matters.
The Riccati recursion
The lowest cost from state is a quadratic, . One step of dynamic programming (pay for this step, then from wherever you land) updates it:
Starting from , each update looks one step further ahead. Once stops changing, it is the infinite-horizon cost and the gain is .
P = Q
repeat:
K = solve(R + B.T @ P @ B, B.T @ P @ A)
P_next = Q + A.T @ P @ (A - B @ K) # the same update, written with K
stop once P_next is (relatively) equal to P
P = P_next
return K computed from the final P
Use a convergence test, not a fixed count: on the cart-pole, 100 iterations still leave about 2 % off.
Tools
cartpole.linearize() returns (A, B) for one 20 ms control period.
The state is (m, m/s, rad, rad/s); the input is the force on the cart (N).
The program
The pole starts about 10° from upright, to either side. The program computes with and , then applies at 50 Hz for 3 s.
Your task
Implement lqr(A, B, Q, R), returning as an array, for any sizes.
The grader compares it with the reference on the cart-pole and three random systems, then checks the pole is within 1° of upright by 2 s, with no request over 10 N and no end-stop hit.