How to design a contact-rich collection rig around the fact that cameras cannot see the moment of contact: sensing, synchronization, and schema.
Why open does not answer the commercial question for robotics datasets, where a dataset's terms really live, and a check to run before training.
Why teleoperation is a collection method and VLA data is a training format, what has to be added between them, and when that work has to happen.
The five axes of egocentric dataset diversity, what the published datasets actually cover, and how to specify coverage before collection starts.
What a reported preference win rate mechanically is, the four disclosures a headline number omits, and how to report one that survives scrutiny.
Why crowd data and expert data capture different signals rather than sitting at two price points, and what that difference changes in your pipeline.
Where teleoperation episodes are lost across collection, annotation, and curation, and how to budget a campaign by yield rather than by session count.
What the measurements show about automated judges on creative output, where a judge earns its place, and where it cannot settle the question at all.
What aesthetic data is, why the published category stays thin where it matters, and the four decisions that determine whether a dataset is usable.
A test that tells you whether your task needs expert judgement or simply a better written specification, before you commit an annotation budget.
Why preference tuning narrows image model output, what the published measurements show, and how to collect preference data that keeps range.
How to separate a real design quality regression from a drifting standard, using a frozen reference set, per-criterion scores, and agreement.
The calibration, timing, and annotation failures that produce a valid but unusable teleoperation dataset, and what to verify before teardown.
The measured ceiling on public human text, why generated data only partly answers it, and the three mechanisms that produce net-new human data.
How to write and audit an annotation specification for a domain nobody on your team understands, using seeded gold items and agreement patterns.
On some tasks the judgment is the label, and no guideline document transfers it. Here is how to tell which tasks those are and how to run expert annotation well.
Crowdsourced annotation fails in ways throughput dashboards are not built to detect. Five failure modes, how to test for each, and where the model stops being appropriate.
The final few percent of cases resists the methods that got you the first 95%, because rarity is a property of the distribution you are sampling from.
Internet video is abundant and free, and it records what happened rather than what was commanded. That missing action channel is the constraint that shapes world model training.
Robot foundation models are not short on trajectories. They are short on diversity, grounded language, failure coverage, and modalities, and more of the wrong data makes them…
Visual realism and physical understanding are measurably different capabilities. Here is what a world model needs in its training data to learn the second one.
Contact is where manipulation succeeds or fails, and it is the signal robot datasets are least likely to contain. Here is what makes it hard to capture and what it costs to fix.
Embodied AI runs on data that has to be produced under a protocol rather than collected from the web, which turns model quality into an operations problem.
Two manipulation datasets of the same size can differ completely in what they teach a policy. Five design decisions, made before collection, account for most of the difference.