Contact Sales

Annotating Without Breaking the Budget: Saving Human Judgment for the Hard Frames

EgoChores is first-person video of real household chores: 27 task categories, head-mounted capture at 1080p and 30fps, un-staged homes, and continuous recordings. Vacuuming, folding laundry, scrubbing a stovetop, cleaning a mirror. It is the work a robot will need to do, recorded from the point of view of the person doing it.

Labeling where the hands are, frame by frame, is slow and expensive work: at 30 frames per second, ten hours of footage holds 2.2 million decisions about where a hand sits and whether it is in shot.

For EgoChores, we built a pipeline where a detector handles the easy frames and human reviewers handle the hard ones. Between two model generations on comparable clips, the share of frames a reviewer had to correct fell from 13.1% to 1.3%. On clips a person checked frame by frame, 99.17% of the boxes in the shipped output sit on the correct hand.

Capture gets you the raw material

Capturing this footage is its own program. Participants across real households wear head-mounted rigs and record ordinary chores in kitchens and bathrooms. Someone has to manage consent, demographic and geographic spread, and the long tail of unusable recordings.

Most labs we talk with don't want to own that work, and the ones who try find it slower and costlier than the model work it feeds. Footage of this kind is scarce, and it is worth buying on its own.

But a vision-language-action (VLA) model doesn't learn from raw footage. It learns from annotations: where the hands are, what they hold, when a task starts and stops.

In One Million Hours of Human Video: Useless Until Someone Labels It, we explained that egocentric video is action-free: it shows what a person did but gives a policy nothing to learn from until someone labels it. That post laid out four annotation layers a usable corpus needs: episode and task segmentation, hand and object geometry, action and primitive labels, and intention-aligned language. This post goes deep on the second layer, hand and object geometry: what it costs at volume and what we built to afford it.

A skilled annotator working a video timeline handles a few thousand decisions an hour, so ten hours of egocentric video takes months of one person's attention. Ten hours also trains nothing on its own. A real training run needs hundreds of hours, and at that scale, labeling every frame by hand is out of reach.

The two obvious workarounds both fail. Hiring more annotators pushes the cost of the dataset beyond what it is worth. Full automation produces a large file of confident guesses that are wrong exactly where manipulation matters most: hands against reflective surfaces, hands holding cloth, hands leaving the frame. Our approach splits the work: models handle the volume, and people handle the judgment.

What failed first: off-the-shelf hand detectors

Free, fast hand detectors already exist, so we started by benchmarking a generally available hand landmarker. We tuned it to its most favorable settings and ran it on held-out footage of mirrors, bathroom tiles, and vacuuming.

The gap comes from the footage itself. General hand models are trained on hands presented to a camera: well lit, centered, close, and still. A head-mounted camera in a real kitchen offers none of that.

Hands swing from window light into the shadow under a counter within a second, and auto-exposure lags the whole way. A hand scrubbing tile blurs across the frame.

Hands also work at the edge of the shot, turned away from the lens, wrapped around a mop handle, or buried in a folded sheet. Skin tones and households vary, and the rig sits at a different height on every participant. Bathroom mirrors add a second pair of hands that belong to no one.

Each chore creates its own detection problem, too. A hand pressing a wadded paper towel against a glass door looks nothing like a hand holding a vacuum, and a general model has seen neither. At 64% recall, an annotator gets a timeline missing a third of its hands, which costs more attention than it saves.

Let the model handle the volume

A detector runs over every frame and proposes a box for each hand. A second model looks at the whole frame, not just the box, and decides whether each hand is visible at all. Neither model is trusted on its own.

Before anything reaches an annotator, a set of filters deletes output the models can't support:

  • Boxes larger than any hand ever recorded in this corpus.
  • Boxes that jump across the frame between consecutive frames, or balloon in place without moving.
  • Boxes that leave a hand, wander, and return to where they started. This usually means the detector locked onto a foot or a mop head for a few frames.
  • Tracks too short to be a hand doing anything.

Each filter has to prove itself against human-checked footage before it ships, and the bar for deleting is deliberately higher than the bar for detecting. A filter that removes one real hand to catch a handful of bad boxes costs a reviewer more than it saves. The filters exist to save reviewer time: a wrong box costs as much as a missing one, because someone still has to notice it, judge it, and fix it.

Send people to the frames that need judgment

In the earlier post we put forth the principle "automate the volume, verify the judgment." The filters cover the volume. For the judgment, annotators receive clips that are already labeled, with low-confidence output removed, so their job becomes reviewing and correcting rather than drawing from scratch.

On easy frames, a second human pass adds nothing. We compared two annotators on the same frames, and their boxes agreed at a median IoU (intersection over union, a measure of box overlap) of 0.775. The detector's boxes had a median IoU of 0.909 against an annotator's. Boxing a clearly visible hand is mechanical work, and the model does it as well as a person.

The hard frames cluster in predictable ways, which matters when you plan an annotation budget. Two thirds of the genuinely ambiguous cases in this corpus involve a soft, light-colored object in the hand, such as a towel, a folded sheet, or a wad of tissue. Under kitchen light, cloth and skin blur into one shape, and deciding where the hand ends takes judgment.

Reflections put hands in the frame that aren't there, and when a hand disappears, someone has to decide whether it left the shot or the camera turned.

A better threshold won't resolve these. They need a person who can see that the bright shape in the hand is a dish towel. Routing clips by model uncertainty puts those frames in front of reviewers and keeps the rest of the timeline out of their way.

Feed the corrections back

Reviewer corrections are the output of the review pass, and they are also training data for the next model generation. Most pre-labeling pipelines stop improving after a generation or two. A pre-filled box that a reviewer accepts without changing teaches the model nothing new: it is the model's own output handed back as supervision. If the training set is full of those accepted boxes, the model mostly learns to reproduce what it already does. We separate corrected frames from accepted ones and weight them accordingly, so the gradient comes from the frames where a person changed something.

The same problem affects measurement. It is the sharpest form of the silent failure we warned about in the earlier post: labels that look plausible but are subtly wrong.

Score a model against annotations that started as its own pre-fills and you measure its agreement with itself. A filter can post an AUC above 0.94 at catching bad boxes on that kind of ground truth and still delete real hands in production, because it learned to agree with the model that generated the labels.

So we hold back clips that no model has touched, annotated from a blank timeline, and no gate or threshold ships until it passes on them. Corrections, clean evaluation sets, and targeted review are what keep the model improving from one generation to the next. Between two model generations on comparable clips, the share of frames a reviewer had to correct fell from 13.1% to 1.3%.

What comes out

These figures come from clips a person checked frame by frame.

Most remaining errors are boxes on the right hand that fit it loosely, the cheapest kind of error to have left. Reviewer effort tells the more useful story: a clip that once needed correction on one frame in eight now needs it on one in eighty.

People still make the judgment calls: deciding whether the bright shape is a towel or a palm, catching the mirror, saying whether a hand left the frame or the model lost it. We just no longer send that person to every frame of every clip.

Where this goes

Hand and object geometry is one of the four annotation layers. Task segmentation, action primitives, and intention-aligned language each have their own version of the 2.2-million-decision problem, and each needs a person somewhere deciding what the pixels don't say. Knowing a corpus's conversion cost before you buy it is most of the work of planning one.

What this looks like as a service

Our services team runs every program in four steps, and EgoChores was no exception:

  1. Define the data and what counts as success.
  2. Design the collection protocol, QA plan, and labeling interface together.
  3. Execute under supervision.
  4. Deliver a verified dataset with its documentation.

The same process applies to warehouse picking, surgical video, in-cabin footage, or whatever domain your policy fails in. You can also use each piece on its own. We can use your footage or source collection.

The free EgoChores sample is one hour of footage captured and annotated for dexterous hand-object manipulation. Request the sample dataset to see our capture and annotation quality for yourself, or talk to the team about scoping your own project.

Related Content