Egocentric Vision
One of the largest unscripted egocentric datasets across industrial and everyday environments. Video, IMU, and audio enriched with 3D hand pose, point tracks, depth, and task annotations.
- 2M+
- Hours
- 50+
- Environments
- 20K+
- Tasks
- >85%
- Hand visibility
This is not a small academic dataset.
One dot = one year of a single camera recording nonstop. 2M+ hours is about 228 of them. Arithmetic on the published total.
See the world where the robot needs to act.
A head-mounted camera keeps hands, tools, objects and contact in frame, from the point of view of the one acting. Hands stay visible more than 85% of the time. Egocentric data is one ingredient of robot learning, not the whole recipe.
Coverage · category share
Top 8 of 26 categories. Midcentury notes these are modeled estimates from sample share, not an audited per-category count.
Difficulty
Specs
- Resolution
- 1920×1080
- Frame rate
- 30 fps
- Field of view
- ~180°
- Clip length
- 3–30 min
- Per participant
- 5–20 h
- Hand visibility
- >85%
- Modalities
- RGB · IMU · audio
- 3D pose
- mm-accurate
- Format
- MCAP / MP4
Files per clip
<id>.v2.mp4Source RGB video, head-mounted capture<id>.v2.imu.jsonSynchronized accelerometer + gyroscope<id>.pose.jsonPer-frame 3D hand keypoints (SLAM-based, mm accuracy)<id>.3d.mp43D pose scene render<id>.2d.mp42D camera-overlay render<id>.v2.jpgPoster frame
Collected with custom head-mounted devices in live work environments. Participants wear the device through normal workflows, typically 5–20 hours each. 11-stage, cost-ordered QA screens every clip before it’s accepted.
Every clip comes from paid contributors who gave explicit consent, with scene-level provenance recorded and non-consenting bystanders blurred during QA before delivery.