Balancing & Optimal Control, step 4
LQR: gains from a cost
Say what you care about as a cost, let the Riccati iteration find the best gain, and tune two controllers: a firm one and a gentle one.
Builds on Feedback: P and PD control, from the free Foundations.
Write and run this step in the simulator with ProSay what you care about
Pole placement works, but which poles do you actually want? The linear-quadratic regulator (LQR) asks a different question: what does a bad future cost? It adds up, over every period from now on,
is a 4×4 weight on the state. A diagonal charges : how much you mind the cart being off centre, moving, the pole leaning, tipping. charges for force. The cheapest controller turns out to be linear again, .
The Riccati iteration
Let be the least cost you can still get from state , the cost-to-go. Start with , which only counts the present, and look one more period ahead each time:
Once stops changing it is the cost over the whole future, and from it is the optimal gain. Stop on convergence, not after a fixed count: on the cart-pole takes hundreds of iterations, and sometimes over a thousand.
Tuning
Only the balance between and matters. Make bigger and force gets dearer: the controller pushes more gently, the pole takes longer to come back, and the cart travels further to catch it. A bigger weight on holds the cart nearer the centre.
Your task
- Implement
lqr(A, B, Q, R)for any number of states and inputs . ReturnK, P: is . - Choose
Q_GENTLEandR_GENTLEfor a gentle controller that never asks for more than 4 N and keeps the cart within 15 cm of the centre.
The program balances from 10° with the firm weights (, ), resets, and balances again with yours. Each must bring the pole within 1° of upright, by 1.5 s (firm) and 2 s (gentle). The plots put the two runs on top of each other.