01The landscape
Two camps, two walls
Manipulation is not the language story. Language is already the most distilled thing humans produce. The physical world is far higher-dimensional and has no corpus.
01Pixel-space world models
They predict frames, not places
Trained to generate what a scene will look like next. Remarkable video, and nothing a body can act on. A plausible next frame is not a metric scene, so nothing in the output tells a robot where anything actually is.
02Vision-language-action
They learn a body, not a task
Trained to map an instruction to joint commands. They work, and the price is tens of thousands of teleoperated demonstrations. Change the robot and the policy does not come with you.
Both camps, and the industry push toward ever more perfect simulators and digital twins, are reaching for a world model that is not currently buildable. The useful thing sits between them, and it is smaller than either: a scene a body can act in, and a record of what the person was trying to do in it.