Notes on the seams: how a bounding box becomes a pose, why the control loop runs at the speed of the slowest honest signal, and what the ROCm training pipeline was actually for.

Explore this article visually3 figures

01

A fresh delivery can contain an old observation

02

Camera, target, hold

Figure 01Data flow

Pixels do not become metres by changing the label

Select an element to explore its role.
Read every explanation
Image plane
The detector gives a location in the image. A bounding-box centre alone leaves the target’s distance from the camera unknown.
Camera frame
Use measured depth, known geometry or a stated plane assumption to recover a 3D point in the camera frame. Check units and calibration provenance.
Robot frame
Apply the calibrated rigid transform, then check reachability and collision constraints. A coordinate transform does not prove a motion is safe.
01 / 03
Image plane

The detector gives a location in the image. A bounding-box centre alone leaves the target’s distance from the camera unknown.

Conceptual coordinate chain. A 2D detection needs depth or explicit geometry before a metric transform is meaningful.

The brief we set ourselves: a creator recording alone says “follow me” or “film the cup”, the arm picks up a small camera from a rest position, points it at the target and keeps the target framed as it moves. Success is judged instantly by anyone watching—is the camera held, is the subject in frame, is the shot steady—which made it a good hackathon task: no metric to argue about.

The stack was a LeRobot-compatible arm, a wrist camera and a scene camera, YOLO for detection over the COCO classes (so “person”, “cup”, “bottle” and friends work out of the box), a small voice-command parser, and a grasp policy trained with imitation learning on episodes we recorded during the event. AMD provided the ROCm GPUs; the training pipeline ran there.

03

A bounding box is not a pose

YOLO gives pixel coordinates in the scene camera's frame. The arm needs a target in its own workspace frame, in metres, with a reachability check. Between the two sits calibration, and the mistake we nearly made was treating it as a setup step done once on the first morning.

Instead we built a small calibration tool that computes the scene-camera-to-arm-base transform from a few taught points and writes the result to a JSON file. The file is loaded at startup and its checksum is logged with every run. When the scene camera was knocked on the second day—which it was—the fix was re-running the tool, not re-tuning anything downstream.

jsoncalibration.json. Boring on purpose: a transform, its provenance, and the workspace limits that gate every planned pose.
{
  "scene_cam_to_base": { "R": [[0.998, -0.052, 0.031], [0.051, 0.998, 0.019], [-0.032, -0.017, 0.999]],
                          "t": [0.412, -0.088, 0.297] },
  "calibrated_at": "2025-12-06T10:41:03Z",
  "residual_mm": 4.2,
  "workspace": { "x": [0.12, 0.48], "y": [-0.30, 0.30], "z": [0.02, 0.35] },
  "camera_grasp_pose": { "approach_offset_m": 0.06, "gripper_close": 0.72 }
}

residual_mm is the one field I would insist on keeping. The residual helps detect a poor fit, but acceptable error depends on the task, geometry and held-out calibration checks. The example values are illustrative, not general acceptance thresholds.

Diagram 02

One loop, three exits

Listenvoice commandTargetYOLO · COCOFramecalibration.jsonPlanjoint limitsExecuteSO-101Verifytarget in frame?10 Hzone observation per turnno detection for 1 s → Pauseoutside reach → Askdrift > 15 % → re-planDashed exits wait for a fresh, valid observation.
The control loop runs at roughly 10 Hz with a fresh observation on every turn. The dashed exits are states of their own: the arm waits there until a new, valid observation arrives, rather than finishing an old plan.

04

The loop runs at the speed of the slowest honest signal

The control loop is simple to state: listen, resolve the target, map it into the workspace, plan a bounded motion, execute one step, verify the target is still where it should be in the frame, repeat. What made it work was refusing to let any stage run ahead of the others. Detection ran at around 10 Hz on the hardware we had; the loop ran at 10 Hz, and one motion step never used a target estimate older than the previous turn.

Three transitions leave the loop, and each is a state rather than an error. No detection for one second: hold position, say so, keep listening. Target outside the workspace box: do not plan, ask the person to move closer. Target drifting more than roughly 15 % of the frame from centre: re-plan from the current observation rather than continuing the previous motion. The temptation in a hackathon is to make the arm keep moving because a moving arm looks impressive. A paused arm that explains why is safer and, in the demo, more convincing.

05

What the ROCm pipeline was for

Figure 03Feedback loop

Demonstration → policy → physical verification

Select an element to explore its role.
Read every explanation
Record
Keep camera observations aligned with robot actions and record the calibration used. Incorrect alignment can teach the wrong response.
Train
Track which episodes and configuration produced the policy. Split episodes deliberately rather than mixing neighbouring frames across sets.
Verify
Test physical outcomes separately from training loss. A low prediction error is not a successful grasp or a stable carrying pose.
Diagnose
Identify whether the failure came from perception, calibration, the policy or verification. Collect another demonstration only when it addresses that failure.
The result informs the next iteration
01 / 04
Record

Keep camera observations aligned with robot actions and record the calibration used. Incorrect alignment can teach the wrong response.

An imitation-learning workflow. Held-out episode evaluation and physical success measure different parts of the system.

The grasp itself—approach the camera on its rest, close the gripper, lift to a carrying pose—was learned rather than scripted, from a few dozen teleoperated episodes recorded in LeRobot's dataset format. The training pipeline on ROCm handled the loop of record, train, evaluate on held-out episodes, publish. We released the dataset and the model artefacts so the result could be reproduced by someone with the same arm.

The point of a pipeline in a two-day event is not scale; it is that when the calibration changed or an episode turned out to be bad, retraining was a command rather than a notebook session. Reproducibility bought us the second day.

06

What broke, concretely

  • “Film the cup” with two cups in view. The parser picked the class; nothing picked the instance. We added a “the closer one” rule and would add pointing next time.
  • The wrist camera was occluded by the gripper for the last 3 cm of the approach. The grasp policy had learned this; the verification step had not, and briefly flagged a lost target on every grasp.
  • Framing drift after a successful grasp: the camera's own weight shifted the wrist estimate by a couple of degrees. Fixed with a re-calibration of the carrying pose after grasp, not before.
  • A voice command during motion. We had no rule; the arm finished the step and then obeyed. That was the right behaviour by luck, and it should have been by design.

None of these were model failures. They were all at seams—between a class and an instance, between a policy and a verifier, between a pose before and after load. That is where the next iteration would spend its time.

07

Separate perception timing from motion control

The roughly 10 Hz loop described here is a task-level perception and planning loop. It should not set the frequency of a motor controller or its safety checks. A 2D bounding box also supplies no metric depth: mapping a target into the arm frame needs depth, known geometry or a stated plane assumption, as well as calibration. Evaluate grasp success, tracking error and lost-target response separately. Calibration residuals describe a fit; they do not by themselves bound collision risk or end-effector error.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read