Principal TPM
What making a robot ad taught us about video models and world knowledge.
Creative Director
Point clouds, calibrated cameras, temporal tracks, and QA belong in one workspace. Here is how the LiDAR interface in Label Studio Enterprise keeps them together.
HumanNet assembles roughly one million hours of egocentric human video, but those million hours are not a dataset. They are raw material, and this post is about the labor that turns raw material into training data.
The synthetic data boom in robotics is real, it's accelerating, and one of the most promising answers to the data bottleneck the field faces. But it has tradeoffs, and it's not a replacement for real-world data.
You fine-tune a Vision-Language-Action model, run it on LIBERO, and post a 95% success rate. The number goes in the paper, the demo video, the investor update. Then you deploy the same model on a real robot in a slightly different room, and it faceplants. Here's why.
Learn why success-only training produces brittle robots, what the new generation of failure-centric datasets looks like, and what a failure-annotation schema contains.
This post explains the data cause behind why language inputs sometimes don't work, where architectural changes fall short, and a fix that is cheap, quantified, and doesn't require collecting a single new trajectory.
Subscribe for news.
A synthesis of the data problem across three years of Vision-Language-Action research, and what it says about where the robotics field goes next.
CEO & Co-Founder