The Idea
The Universal Manipulation Interface line of work made a simple trade explicit. Instead of driving a robot through a task, the demonstrator holds a gripper, one that matches the robot's end effector, and performs the task by hand. A wrist-mounted camera rides on the gripper. From the video and the device's tracking, the pipeline recovers the end-effector trajectory and gripper aperture over time. The output is a demonstration in end-effector space: not tied to any arm, replayable on any robot whose controller accepts end-effector targets.
Decoupling the demonstration from the robot removes the two bottlenecks that cap teleoperation throughput: robot time and rig latency. A demonstrator with a handheld gripper works at nearly natural speed, several collectors work in parallel without several robots, and capture can happen in scenes no lab robot will ever visit.
What a UMI-Style Episode Contains
- Egocentric wrist video from the gripper-mounted camera, the same viewpoint the deployed policy will see from its own wrist camera.
- Recovered end-effector trajectory, the 6-DoF pose track over the episode, produced by visual-inertial tracking and validated per episode.
- Gripper state, aperture over time, aligned to the trajectory.
- Task and stage labels against the spec, plus device calibration and a data card, the same delivery discipline as every collection we run.
What it does not contain is on-robot proprioception and force. The trajectory is recovered, not commanded, so precision has a floor set by tracking quality, and contact-rich force signatures are absent. Tasks that live and die on force control belong on a teleop rig instead.
Failure Modes and the QA That Catches Them
Handheld capture has known ways to go wrong, and an operation that runs it at volume is defined by how it handles them:
- Tracking loss. Fast motion and textureless scenes break pose recovery. Every episode's trajectory is checked for continuity and drift, and episodes with unrecoverable segments are re-collected rather than interpolated.
- Motion blur and exposure. Natural human speed is the method's selling point and its camera's enemy. Studio lighting, shutter discipline, and per-clip image QA keep the video trainable.
- Kinematic infeasibility. Human wrists reach poses robot arms cannot. Demonstrator training plus post-hoc feasibility filtering against the target embodiment keeps delivered episodes executable.
- Framing drift. The action must stay in the wrist camera's view. Collectors are trained on framing, and clips are gated on it.
This is exactly the kind of production discipline that makes capture as a service sensible to buy: the method is public, the reliable operation of it is the product. Ours runs in purpose-built capture studios in Asia Pacific, with mobile capture available when a spec calls for in-the-wild scene diversity.
What to Pin Down in the Spec
Three decisions shape a handheld collection more than any other. First, the end effector: the handheld gripper should match the deployment gripper's geometry and stroke, because finger shape changes which grasps the demonstrations contain. Second, the scene policy: studio-controlled scenes maximize episode validity and repeatability, while in-the-wild capture maximizes visual diversity, and the right mix depends on whether your bottleneck is precision or generalization. Third, the labeling depth: trajectory plus success labels is the floor; stage segmentation and grasp annotations cost little at collection time and open up curriculum and subgoal training later. Each of these is one line in the collection spec, and our spec checklist covers the rest.
Where It Fits in a Training Mix
The pattern that keeps winning is layered: a broad UMI-style corpus for scale and diversity, co-trained or followed by a focused on-robot teleop set that grounds the policy in the target embodiment's real dynamics. Buyers pursuing a single fixed platform sometimes skip the handheld layer entirely, and buyers who have not settled on hardware sometimes collect only in end-effector space so the corpus survives a platform change. Both are legitimate specs; what matters is choosing deliberately. Bare-hand egocentric video can sit underneath both as pretraining signal. How to split a data budget across the three layers is a spec question, and our spec guide walks through it.
Tell us the task, the target end effector, and the scene diversity you need. We reply with a capture plan, an episode estimate, and a 60-day delivery date. Get a collection quote.

