Notes on the seams: how a bounding box becomes a pose, why the control loop runs at the speed of the slowest honest signal, and what the ROCm training pipeline was actually for.
Explore this article visually3 figures
01
A fresh delivery can contain an old observation
02
Camera, target, hold
Figure 01Data flow
Pixels do not become metres by changing the label
Read every explanation
- Image plane
- The detector gives a location in the image. A bounding-box centre alone leaves the target’s distance from the camera unknown.
- Camera frame
- Use measured depth, known geometry or a stated plane assumption to recover a 3D point in the camera frame. Check units and calibration provenance.
- Robot frame
- Apply the calibrated rigid transform, then check reachability and collision constraints. A coordinate transform does not prove a motion is safe.
The detector gives a location in the image. A bounding-box centre alone leaves the target’s distance from the camera unknown.
The brief we set ourselves: a creator recording alone says “follow me” or “film the cup”, the arm picks up a small camera from a rest position, points it at the target and keeps the target framed as it moves. Success is judged instantly by anyone watching—is the camera held, is the subject in frame, is the shot steady—which made it a good hackathon task: no metric to argue about.
The stack was a LeRobot-compatible arm, a wrist camera and a scene camera, YOLO for detection over the COCO classes (so “person”, “cup”, “bottle” and friends work out of the box), a small voice-command parser, and a grasp policy trained with imitation learning on episodes we recorded during the event. AMD provided the ROCm GPUs; the training pipeline ran there.
03
A bounding box is not a pose
YOLO gives pixel coordinates in the scene camera's frame. The arm needs a target in its own workspace frame, in metres, with a reachability check. Between the two sits calibration, and the mistake we nearly made was treating it as a setup step done once on the first morning.
Instead we built a small calibration tool that computes the scene-camera-to-arm-base transform from a few taught points and writes the result to a JSON file. The file is loaded at startup and its checksum is logged with every run. When the scene camera was knocked on the second day—which it was—the fix was re-running the tool, not re-tuning anything downstream.
{
"scene_cam_to_base": { "R": [[0.998, -0.052, 0.031], [0.051, 0.998, 0.019], [-0.032, -0.017, 0.999]],
"t": [0.412, -0.088, 0.297] },
"calibrated_at": "2025-12-06T10:41:03Z",
"residual_mm": 4.2,
"workspace": { "x": [0.12, 0.48], "y": [-0.30, 0.30], "z": [0.02, 0.35] },
"camera_grasp_pose": { "approach_offset_m": 0.06, "gripper_close": 0.72 }
}residual_mm is the one field I would insist on keeping. The residual helps detect a poor fit, but acceptable error depends on the task, geometry and held-out calibration checks. The example values are illustrative, not general acceptance thresholds.
Diagram 02
One loop, three exits
04
The loop runs at the speed of the slowest honest signal
The control loop is simple to state: listen, resolve the target, map it into the workspace, plan a bounded motion, execute one step, verify the target is still where it should be in the frame, repeat. What made it work was refusing to let any stage run ahead of the others. Detection ran at around 10 Hz on the hardware we had; the loop ran at 10 Hz, and one motion step never used a target estimate older than the previous turn.
Three transitions leave the loop, and each is a state rather than an error. No detection for one second: hold position, say so, keep listening. Target outside the workspace box: do not plan, ask the person to move closer. Target drifting more than roughly 15 % of the frame from centre: re-plan from the current observation rather than continuing the previous motion. The temptation in a hackathon is to make the arm keep moving because a moving arm looks impressive. A paused arm that explains why is safer and, in the demo, more convincing.
05
What the ROCm pipeline was for
Figure 03Feedback loop
Demonstration → policy → physical verification
Read every explanation
- Record
- Keep camera observations aligned with robot actions and record the calibration used. Incorrect alignment can teach the wrong response.
- Train
- Track which episodes and configuration produced the policy. Split episodes deliberately rather than mixing neighbouring frames across sets.
- Verify
- Test physical outcomes separately from training loss. A low prediction error is not a successful grasp or a stable carrying pose.
- Diagnose
- Identify whether the failure came from perception, calibration, the policy or verification. Collect another demonstration only when it addresses that failure.
Keep camera observations aligned with robot actions and record the calibration used. Incorrect alignment can teach the wrong response.
The grasp itself—approach the camera on its rest, close the gripper, lift to a carrying pose—was learned rather than scripted, from a few dozen teleoperated episodes recorded in LeRobot's dataset format. The training pipeline on ROCm handled the loop of record, train, evaluate on held-out episodes, publish. We released the dataset and the model artefacts so the result could be reproduced by someone with the same arm.
The point of a pipeline in a two-day event is not scale; it is that when the calibration changed or an episode turned out to be bad, retraining was a command rather than a notebook session. Reproducibility bought us the second day.
06
What broke, concretely
- “Film the cup” with two cups in view. The parser picked the class; nothing picked the instance. We added a “the closer one” rule and would add pointing next time.
- The wrist camera was occluded by the gripper for the last 3 cm of the approach. The grasp policy had learned this; the verification step had not, and briefly flagged a lost target on every grasp.
- Framing drift after a successful grasp: the camera's own weight shifted the wrist estimate by a couple of degrees. Fixed with a re-calibration of the carrying pose after grasp, not before.
- A voice command during motion. We had no rule; the arm finished the step and then obeyed. That was the right behaviour by luck, and it should have been by design.
None of these were model failures. They were all at seams—between a class and an instance, between a policy and a verifier, between a pose before and after load. That is where the next iteration would spend its time.
07
Separate perception timing from motion control
The roughly 10 Hz loop described here is a task-level perception and planning loop. It should not set the frequency of a motor controller or its safety checks. A 2D bounding box also supplies no metric depth: mapping a target into the arm frame needs depth, known geometry or a stated plane assumption, as well as calibration. Evaluate grasp success, tracking error and lost-target response separately. Calibration residuals describe a fit; they do not by themselves bound collision risk or end-effector error.
More articles
09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read