Guide · 2026

Robot training data from hospitality work: what labs need, and what a clean pilot looks like.

By Lilian Boboc, founder of Hospara · Updated 2026-09-30 · Based on a Bali operation of approximately 600 units

Humanoid and manipulation models learn from first-person video of people doing real tasks. Housekeeping, kitchen and laundry work is exactly the kind of repeated, procedural manipulation those models need. This guide sets out what a lab-grade capture program has to get right, based on the pilot Hospara is preparing for Q4 2026.

Direct answer

First-person (egocentric) video of housekeeping, kitchen, laundry, pool and garden tasks, captured to a specification agreed with a robotics lab, with a consent chain covering the property, the worker and any incidental capture. It is used to train manipulation and humanoid models. Hospara is preparing a Q4 2026 pilot in Bali.

What egocentric video is

Video recorded from the worker's point of view, usually with a head-mounted phone or camera, showing hands, tools and objects during a task. Labs use it to pre-train and fine-tune models that map perception to action. It is video only unless extra sensors are added; it does not contain robot actions, joint states or depth by itself.

Why hospitality tasks

Bed making, bathroom reset, mopping, dishwashing, folding, pool skimming and garden work are repeated hundreds of times a week, follow written procedures, and involve deformable objects, liquids, tools and varied lighting. That combination is scarce in existing datasets, which skew to kitchens in a few countries.

What a specification contains

Task family, repetitions, camera device and mount, angle, resolution and frame rate, pace, left and right handed operators, failure and recovery cases, lighting conditions, environment variety, and the acceptance criteria the buyer's QA will apply. Capture without a specification produces footage nobody can use.

The consent chain

A separate, specific authorisation from each property owner; a signed participation agreement per worker in their own language; recording only in vacant units; a capture-prevention policy for guests, documents and screens; blur before export; a per-clip consent log; and a written quarantine and deletion procedure for accidental capture. Under Indonesia's PDP Law, lawful basis and cross-border transfer safeguards are reviewed with counsel before anything leaves the country.

Delivery

Task-segmented clips with timestamps and JSON manifests, blur applied, consent log per clip. Packaging in RLDS or LeRobot directory layouts on request, carrying video and manifests only. Annotation of hands or objects through partners on request.

How Hospara is approaching it

A Q4 2026 pilot in participating Bali properties, starting with a paid calibration stage in one or two authorised units, then volume agreed per specification. Nothing has been recorded or delivered yet. Details on the labs page.

Related

Questions

Straight answers.

Does egocentric video include robot actions or depth?
No. It is video plus task manifests unless additional sensors are agreed. Packaging in RLDS or LeRobot layouts does not add actions, state or depth.
How is worker consent handled?
Each worker signs a participation agreement in their own language, is paid per accepted hour on top of salary, and can stop future recording at any time without any effect on their employment. Training and recording time count as working hours.
Can a lab ship its own capture hardware?
Yes. The default is a head-mounted phone; GoPro, smart glasses or a lab rig can be used with the lab's calibration procedure.