The Service, in One Sentence
You describe a task, an embodiment, and a training pipeline; the service returns a corpus of demonstration episodes, collected by trained operators against a written spec, checked episode by episode, and delivered in the format your training code already reads. Everything below is that sentence expanded into the six stages every serious engagement passes through.
The buyers are recognizable: robotics teams whose engineers are spending policy-training time running capture sessions, foundation model groups that need coverage on tasks and objects the open corpora never touch, and product teams that need one embodiment doing one job reliably. What they have in common is that data collection is on their critical path but is not their business. The service exists to take it off the path.
Stage One: The Spec Is the Product
A collection spec is a contract about physics. It fixes the task and its success criteria, the action space your controller consumes, camera count and placement, what varies between episodes and what stays pinned, and how failures and recoveries are represented. Most disappointing datasets were doomed at this stage, not on the collection floor: an action space mismatch or an ambiguous success criterion survives every later quality check because the data is faithfully wrong. We publish the exact checklist we use in the spec guide, and the first working session of any engagement is spent arguing about it, on purpose.
Stage Two: A Pilot Batch, While Changes Are Cheap
Before full collection, a pilot batch of real episodes lands in your training pipeline within two weeks. You load it, train on it if you want, and inspect cameras, alignment, labels, and format against your own tooling. Camera placement problems, ambiguous edge cases, and format friction all surface here, where fixing them costs a revision to the spec instead of a re-collection of the corpus. Full collection starts only after you sign off on pilot episodes.
Stage Three: Trained Operators, Not Crowd Labor
Demonstration data is human skill made machine-readable, so operator quality is dataset quality. Our operators train on each task until their success rate and trajectory smoothness stabilize before a production episode is recorded, and the protocol scripts what varies between episodes: object placement, distractors, lighting, initial states. Recovery demonstrations are staged deliberately, because a corpus of clean successes teaches a policy that has never seen trouble. The capture method follows the task rather than the other way around: teleoperated rigs where on-robot state-action pairs are required, UMI-style handheld grippers where volume and scene diversity matter more, and instrumented egocentric capture where full-speed human dexterity is the signal.
Stage Four: Purpose-Built Studios
Collection runs in purpose-built capture studios in Asia Pacific, and the word purpose-built is doing real work. Fixed camera geometry, controlled lighting, maintained object and material libraries, and engineered reset stations are what make two collection days six weeks apart statistically compatible, and what make throughput a schedule instead of a hope. Resets get jigs and two-station alternation so rigs never idle while a scene is rebuilt; this unglamorous logistics layer is most of the difference between a service and a lab intern with a camera.
Stage Five: QC on Every Episode
Batch-level sampling misses exactly the defects that poison training runs, so checks run per episode: timestamp monotonicity and cross-stream synchronization, dropped frame detection, calibration presence, action-observation alignment, and label consistency against the written success criteria. Episodes that fail are re-collected, never patched, and the check results ship with the data so your team can audit rather than trust.
Stage Six: Delivery in the Format You Train In
The corpus arrives in LeRobot, RLDS, or GR00T-compatible form, with calibration files, QC reports, and a loader script that reads the dataset unmodified. If the format question is still open on your side, the formats guide lays out the trade-offs for buyers. Standard scopes deliver within 60 days of the signed spec; unusual scopes get an honest timeline up front instead of a promise. Custom collections are exclusive to the buyer and delivered under full ownership.
When to Buy This Instead of Building It
Building an internal collection operation makes sense when data collection is your product or when you will collect continuously for years. It rarely makes sense for a team that needs one focused corpus: the rigs, the operator training curve, the reset engineering, and the QC tooling are all fixed costs the first dataset has to absorb alone. The buying decision, including when open corpora are genuinely enough, is covered across the full training data service and its buyer guides. What a service cannot do is rescue a vague goal; if you cannot state the success criterion for the task, no volume of episodes will state it for you.
Tell us the task, the embodiment, and the format you train in. We come back with a collection spec draft, a pilot plan, and a delivery date. Get a collection quote.

