HumanNet assembles roughly one million hours of egocentric human video, but those million hours are not a dataset. They are raw material, and this post is about the labor that turns raw material into training data.
The synthetic data boom in robotics is real, it's accelerating, and one of the most promising answers to the data bottleneck the field faces. But it has tradeoffs, and it's not a replacement for real-world data.
You fine-tune a Vision-Language-Action model, run it on LIBERO, and post a 95% success rate. The number goes in the paper, the demo video, the investor update. Then you deploy the same model on a real robot in a slightly different room, and it faceplants. Here's why.
Learn why success-only training produces brittle robots, what the new generation of failure-centric datasets looks like, and what a failure-annotation schema contains.
This post explains the data cause behind why language inputs sometimes don't work, where architectural changes fall short, and a fix that is cheap, quantified, and doesn't require collecting a single new trajectory.
A synthesis of the data problem across three years of Vision-Language-Action research, and what it says about where the robotics field goes next.
CEO & Co-Founder