Bazyx

Egocentric world models

Make robotsunderstand humans

Do the task once with your own hands. The robot learns it from whatever cameras were watching — no teleoperation, no scripted path.

Visualisation · procedural scene

The landscape

Two camps, two walls

Manipulation is not the language story. Language is already the most distilled thing humans produce. The physical world is far higher-dimensional and has no corpus.

01Pixel-space world models

They predict frames, not places

Trained to generate what a scene will look like next. Remarkable video, and nothing a body can act on. A plausible next frame is not a metric scene, so nothing in the output tells a robot where anything actually is.

02Vision-language-action

They learn a body, not a task

Trained to map an instruction to joint commands. They work, and the price is tens of thousands of teleoperated demonstrations. Change the robot and the policy does not come with you.

Both camps, and the industry push toward ever more perfect simulators and digital twins, are reaching for a world model that is not currently buildable. The useful thing sits between them, and it is smaller than either: a scene a body can act in, and a record of what the person was trying to do in it.

The approach

A canonical scene, and the semantics on top of it

Three layers. A foundation model over the top two, and a bottom layer that makes it portable across bodies.

  1. Layer 03

    Intent and task semantics

    What the person was trying to do, in what order, and why that grip and not another one. The part a demonstration actually carries and a trajectory does not.

  2. Layer 02

    Canonical metric 3D scene

    One coherent, measurable scene fused from however many cameras happen to be watching — a headset, the robot’s own wrist, a phone on a stand — with poses changing online as people and robots move. Geometry a body can act in, rather than pixels from a viewpoint.

  3. Layer 01

    Hardware abstraction

    Embodiment, actuators, camera intrinsics and placement. Different bodies resolve into the same scene, which is what lets a skill move between them.

Why now

The backbones only just landed

Self-supervised visual features, feed-forward 3D from unposed images, and video-native prediction all became usable across 2024 and 2025. None of this was buildable two years ago.

Built on

  • DINOv3 · 2025
  • VGGT · 2025
  • V-JEPA 2 · 2025

Public backbones, with our own forecaster on top — the model that predicts how the scene and the hands move next, which is what turns a reconstruction into something a robot can act ahead of.

Where it runs
Households, machine shops, construction sites. The shared requirement is a robot and a person in the same space on the same task, with the person able to correct it.
First users
Robot-learning labs and student teams, who already have the robot and the task and are missing the data.

Status

Runs today
The forecaster, and the egocentric reconstruction pipeline it runs on: an arbitrary set of camera views fused into one metric scene, poses solved online.
Next
The same loop end to end on a real task — human video in, robot motion out. That demo is what the current work is pointed at.

The loop

Teach it once, then work alongside it

Closer to how you work with a capable colleague than to how you program a machine.

  1. 01

    Demonstrate

    A person does the task once with their own hands, wearing camera glasses. Any other view in the room joins in — the robot’s wrist camera, a phone on a stand. No teleoperation rig, no scripted trajectory, no motion capture suit.

  2. 02

    Ask

    The robot asks when it is unsure, instead of failing confidently. Uncertainty becomes a question rather than a dropped part.

  3. 03

    Co-work

    Human and robot in the same space on the same task, with the person able to correct it mid-task rather than after it.

  4. 04

    Autonomy

    Per-environment fine-tuning runs online until the correction loop goes quiet. Full automation is the end state, not the starting requirement.

Where this goes

What we are building toward

Everything in this section is where the work is aimed, not what runs today.

A foundation model for intent

Human intention and task semantics, transferred to robots. Collaboration first, because that is what is reachable now, with full autonomy as the end state rather than the entry price.

Simulation built from real scenes

Environments reconstructed from point clouds of real, unknown spaces, more dynamic than what current simulators offer. Articulation data is the gap, and it is the one we have to close.

Portability and long horizons

Easier cross-embodiment, better handling of long-horizon tasks, and policies that keep improving online in the environment they were actually deployed in.

Get in touch

We are looking for thefirst people to build on it.

If you run a robot-learning lab, build with manipulation, or invest in this, we should talk.