Start With the Free Data
The open robot datasets are one of the best things that ever happened to this field, and any custom collection pitch that pretends otherwise is selling you something. Before you spend a dollar on bespoke data, you should know exactly what the open corpora give you. Often the right first move is to train on them and see where your policy breaks. The break points are your custom data spec.
What the Open Corpora Contain
- DROID. Roughly 76,000 teleoperated episodes collected across hundreds of real scenes on a standardized Franka arm setup with multi-view cameras. Strong scene diversity by robot-dataset standards, tabletop manipulation focus, one embodiment.
- AgiBot World. Over a million trajectories collected on a fleet of standardized humanoid platforms in staged home, retail, and industrial settings, including some long-horizon and bimanual work. Serious scale, one platform family, staged environments.
- Open X-Embodiment. The aggregation effort: over a million episodes pooled from dozens of labs and more than twenty embodiments in RLDS form. Unmatched embodiment breadth, with the heterogeneity that pooling implies: mixed camera setups, action conventions, and label quality.
All three are pretraining gold. Policies co-trained on these mixtures show measurably better generalization than policies trained from scratch, and every serious manipulation stack should sit on some open prior.
When Open Data Is Enough
- Your task is common tabletop manipulation, pick, place, wipe, open, close, on ordinary rigid objects.
- Your embodiment matches or closely resembles a well-covered platform, a Franka arm being the clearest case.
- You are doing representation learning, pretraining, or method research, where coverage breadth matters more than task fidelity.
- You can tolerate a residual gap and close it with a modest amount of your own on-robot fine-tuning data.
If all four hold, stop here. You do not need us yet, and the money is better spent on compute.
When Open Data Runs Out
- Task coverage. The corpora concentrate on rigid tabletop work. Deformables thin out fast, and skilled operations, sewing, ironing, fastening, dressing, are essentially absent. If your roadmap lives there, the data does not exist to download; see our deformable and garment pages.
- Embodiment. Your camera placement, gripper, and action space differ from the corpus norm, and every difference is a transfer tax your policy pays at deployment.
- Domain. Your objects, materials, and scenes, the reflective, the transparent, the industrial, are not in staged home kitchens.
- Labels and criteria. Open success labels follow the collectors' definitions, not yours. Evaluation-grade data needs your criteria applied consistently.
- Exclusivity. Everyone trains on the open corpora, including your competitors. Data no one else has is one of the few durable advantages left in a world of shared architectures.
The Pattern That Works: Open Prior, Custom Core
The choice is not either-or. The recipe that keeps producing deployed policies is layered: pretrain or co-train on open mixtures for the general manipulation prior, then collect a focused custom set covering your task, your embodiment, and your failure modes. The custom layer is smaller than a from-scratch estimate would suggest, precisely because the open prior does the generic work. A pilot batch trained into your pipeline is the honest way to size it, which is why every engagement we run starts with one.
| Dimension | Open Corpora | Custom Collection |
|---|---|---|
| Cost | Free | Paid, scoped per collection |
| Task match | Whatever was collected | Your task, to spec |
| Embodiment match | Corpus platforms | Your action space and cameras |
| Deformables and skilled work | Thin to absent | Our specialty |
| Success criteria | Collectors' definitions | Yours, written and audited |
| Exclusivity | None, by design | Full buyer ownership |
| Best role | Pretraining prior | Task competence and evaluation |
Common Questions
Should I train on open data before buying custom data?
Usually yes. Pretraining or co-training on open corpora gives you a general manipulation prior at zero data cost. Custom collection then covers your task, your objects, and your embodiment where open coverage ends.
Is open data ever enough on its own?
For common tabletop tasks on a widely used arm, with tolerance for a sim-to-real style gap, it can be. Teams working on standard pick and place with a Franka arm are the best served by DROID. The further your task is from that center, the thinner the coverage gets.
How much custom data do I actually need if I co-train?
Less than a from-scratch estimate suggests. A focused custom set covering your task's states and failure modes, co-trained with an open prior, is the pattern most current fine-tuning recipes assume. Your pilot batch is the honest way to size it.
Describe your task and we will tell you plainly if open data covers it. If it does not, you get a scoped custom collection plan with a 60-day delivery date. Get a collection quote.

