Step 1
The pinhole camera
Predict where a 3-D point shows up in the camera image, using the camera's pose and its intrinsics.
From positions to pixels
In Pick & Place, world.balls() handed you exact positions. Real robots have to see. A camera now stands at the far end of the table, facing the robot. camera.capture() returns an Image whose rgb is a 240 × 320 × 3 array of bytes. Pixel is column and row , so its colour is img.rgb[v, u].
Before a robot can find things in a picture, it has to know where things would appear. That's the pinhole model, and it has two steps.
1. Into the camera's frame
camera.pose is a 4 × 4 transform from camera to world coordinates. Its rotation = T[:3, :3] holds the camera's axes as columns, and = T[:3, 3] is the camera's position. Inverting it moves a world point into the camera frame:
| Camera axis | Points |
|---|---|
| right in the image | |
| down in the image | |
| forward, out of the lens |
2. Onto the image
The intrinsic matrix (camera.K) scales by the focal length , measured in pixels, and shifts by the principal point :
Dividing by the depth gives and . That division is why distant things look small.
Your task
Implement project(p) so it returns np.array([u, v]). The starter captures an image and draws a white cross wherever project says each ball is. Open the Images tab: with a correct project, every cross sits on its ball.
The grader compares project with the real camera on 20 random points. Each must be within 0.5 px.