A first-person human hand picks up a steel bracket and places it in a parts bin. The frame freezes and separates into RGB, depth, 3D hand pose, point tracks, IMU motion and task annotation layers. One clip becomes a field of thousands. A robot learns the same task, succeeds once, then fails as position, lighting and orientation change. The scene becomes a digital twin in Midcentury Matrix and multiplies into thousands of parallel scenarios. One fails; its replay becomes a new training scenario; the robot succeeds in simulation, then back in the real world. Conceptual visualization.

Conceptual visualizationIllustration of a head-mounted, first-person capture
Skip film ↓

Data and Simulation Infrastructure for Physical AI

Real-world egocentric data and high-fidelity simulation at scale for training, evaluating, and deploying robots.

Real→Data→Train→Simulate→Fail→Learn→Real

How to read this page

  • Current product Offered today
  • Early access Available by request
  • Research Published research
  • Coming soon Announced, nothing published yet
  • Demonstration Interactive toy in your browser
  • Conceptual visualization Illustration, not real footage

01Mission · the bottleneck

A demo is one scenario. Reality is every scenario.

One robot. Controlled environment. Perfect conditions.Conceptual visualization
Different objects, surfaces, lighting and positions. The unexpected.Conceptual visualization

Physical AI produces impressive demos. Yet very few robots ever reach the real world.

02Datasets

High-fidelity data for the real world

Large-scale robotics and simulation data, paired with rich action supervision and dense annotations for the next generation of physical AI.

Current product

Egocentric Vision

One of the largest unscripted egocentric datasets across industrial and everyday environments. Video, IMU, and audio enriched with 3D hand pose, point tracks, depth, and task annotations.

2M+
Hours
50+
Environments
20K+
Tasks
>85%
Hand visibility
BaseLayers

This is not a small academic dataset.

1 hour1 clip-hour
1 day24 hours
1 year8,760 hours
100 years876,000 hours
2M+ hours≈ 228 years of continuous footage

One dot = one year of a single camera recording nonstop. 2M+ hours is about 228 of them. Arithmetic on the published total.

See the world where the robot needs to act.

FIXED CAMERAHANDS SMALL · PARTLY OCCLUDED
Third person: the camera sees the person.
First person: the camera sees the task.

A head-mounted camera keeps hands, tools, objects and contact in frame, from the point of view of the one acting. Hands stay visible more than 85% of the time. Egocentric data is one ingredient of robot learning, not the whole recipe.

Coverage · category share

  • Food & beverage14.5%
  • Mechanical & small-parts assembly10%
  • Construction9.5%
  • Cleaning8.2%
  • Lifestyle / home7.9%
  • Hospitality6.2%
  • Repair services6%
  • Fashion5.9%

Top 8 of 26 categories. Midcentury notes these are modeled estimates from sample share, not an audited per-category count.

Difficulty

Easy 6.7%Medium 64.4%Hard 28.9%

Specs

Resolution
1920×1080
Frame rate
30 fps
Field of view
~180°
Clip length
3–30 min
Per participant
5–20 h
Hand visibility
>85%
Modalities
RGB · IMU · audio
3D pose
mm-accurate
Format
MCAP / MP4

Files per clip

  • <id>.v2.mp4Source RGB video, head-mounted capture
  • <id>.v2.imu.jsonSynchronized accelerometer + gyroscope
  • <id>.pose.jsonPer-frame 3D hand keypoints (SLAM-based, mm accuracy)
  • <id>.3d.mp43D pose scene render
  • <id>.2d.mp42D camera-overlay render
  • <id>.v2.jpgPoster frame

Collected with custom head-mounted devices in live work environments. Participants wear the device through normal workflows, typically 5–20 hours each. 11-stage, cost-ordered QA screens every clip before it’s accepted.

Every clip comes from paid contributors who gave explicit consent, with scene-level provenance recorded and non-consenting bystanders blurred during QA before delivery.

Current product · built to spec

Custom Gameplay Environments

When you control the world, you control the data.

Custom, on-demand gameplay data from environments built to spec — any camera angle from first-person to top-down — with frame-aligned inputs, telemetry, camera state, and engine G-buffers.

50K+
Hours supported
100+
Environment types
Engine-level
Signals
G-bufferStreams
Resolution
up to 4K
Frame rate
up to 60 fps
Formats
MP4 / EXR / JSONL
Engines
Unreal · Unity
Core signals
RGB · inputs · camera · telemetry
Engine outputs
Depth · normals · albedo · motion

Capability library · representative slice

12
Environment families
48
Action types
8
Camera views
10
Engine-level outputs

Sample pilot spec: 100 environments · 10,000 pilot hours.

03Simulation

Midcentury Matrix

Agentic simulation platform to design, test, and train physical AI across massively parallel cloud environments.

  1. 01 DesignDigital twins from real deployment conditions, combining classical simulation with learned physics.
  2. 02 TestMassively parallel GPU evaluation across thousands of scenarios. Replay failures, catch regressions before hardware.
  3. 03 TrainTurn failures into new scenarios and training experience. Feed real-world data back into policies.
Demonstration

“Digital twins from real deployment conditions, combining classical simulation with learned physics.”

Scenario distribution (toy)

Drag the divider. Ghost outlines are 28 sampled part poses. These sliders drive this page’s toy model, not Matrix’s interface.

—The loop

Real-world data scales the simulation. Simulation turns failure into experience.

Midcentury builds both halves of the loop, so each can feed the other. Midcentury describes Matrix as a frontier simulation platform scaled with our real-world data to evaluate and post-train policies at scale.

  1. 01Real worldPeople work in real environments.
  2. 02Human actionFirst-person capture of the task as it happens.
  3. 03DataVideo, IMU, audio → 3D hand pose, point tracks, depth, task labels.
  4. 04ModelHuman action data for pretraining physical AI.
  5. 05SimulationDigital twins of real deployment conditions.
  6. 06FailureMassively parallel evaluation surfaces what breaks.
  7. 07TrainingFailures become new scenarios and training experience.
  8. 08Better policyReal-world data fed back into policies. Then back to reality.
DATAMATRIXREAL WORLDHUMAN ACTIONDATAMODELSIMULATIONFAILURETRAININGBETTER POLICYTHE ACTION

04Research

Our contributions to physical intelligence

See all research
ProjectProblemContributionStatus

MC-EgoHands

A frontier 3D harness for human egocentric video.

Illustration. See the report for Midcentury’s real reconstructions.

Egocentric video shows skilled human work. Robot learning needs that work as structured 3D motion it can review, retarget and train on.

  • Reconstructs both hands and the camera, in a shared world frame
  • Inputs: MP4, IMU JSON and MCAP · monocular and calibrated stereo
  • Source adapters: MC-EgoHands, EgoDex and HoloAssist
  • Export: LeRobot v3 episodes · robot targets for Panda / LIBERO and Yam
  • Annotations used in VLA and WAM projects
Published hand-pose results (PA-MPJPE, mm)
ARCTIC3.6665,000 camera images
HOT3D3.9765,000 camera images
H2O5.2409,998 hand instances
FreiHAND4.3025,000 camera images

PA-MPJPE in mm, lower is better. Custom evaluations, not official leaderboard submissions or matched comparisons with other methods; they do not establish unseen-dataset generalization. Published by Midcentury as a research preview.

Research previewRead the reportJoin the waitlist →

MC-Shade

Real-time photorealistic rendering for any simulation.

Not yet published. No architecture, performance or quality figures have been released.

Coming soon

MC-PhysBench

The first long-horizon physics benchmark for world models.

Not yet published. No benchmark tasks or results have been released.

Coming soon

06Company

Built by people who’ve worked at the frontier of robotics data and models.

Team members’ backgrounds include Stanford AI Lab, OpenAI, DeepMind, NVIDIA, Scale AI and Invisible, with prior work connected to RoboNet, GDPVal and NVIDIA Cosmos 3.

  • Stanford AI Lab
  • OpenAI
  • DeepMind
  • NVIDIA
  • Scale AI
  • Invisible