Orply.

World Models Could Ease Robotics’ Training-Environment Bottleneck

Fei-Fei LiEd LudlowBloomberg TechnologyTuesday, September 22, 20265 min read

World Labs CEO Fei-Fei Li argues that spatial intelligence—models that generate controllable, geometrically structured 3D environments—could extend AI beyond language and provide needed training and evaluation settings for robotics. She says the company’s Atlas model is designed to reconstruct and generate navigable worlds from images, while maintaining camera control and visual consistency. Li also argues that safety evaluation cannot rest with AI companies alone, calling for independent benchmarks and participation from academia, government, industry and civil society.

Atlas is meant to make generated worlds navigable and reconstructable

Fei-Fei Li distinguishes World Labs’ world models from language models by what they produce. Language models output words, which can become code or conversation; the systems she describes generate pixels alongside the geometric structure of a three- or four-dimensional environment. The goal is not simply to make an image, but to produce a world that creators, designers and robotics developers can inspect and control.

What we generate are not words. The models generate beautiful pixels, very accurate geometric structure of the 3D or 4D world.

Fei-Fei Li · Source

Atlas, World Labs’ latest model, is intended to combine image generation with intricate 3D reconstruction. Li says it allows camera control and pixel consistency with the real or generated 3D world, addressing what she calls a hard computer-vision problem: producing visually compelling output while preserving detailed spatial structure. Asked whether Atlas solves that problem, she said it does. Asked how it compares with specialized systems, she said it outperforms state-of-the-art specialized models on those tasks.

A World Labs demonstration showed the intended workflow. An input image of a castle exterior or church interior appeared beside a camera path and generated output views through a reconstructed scene. The on-screen labels—“Input Image,” “Camera Path” and “Output”—made the proposition concrete: given an image and a requested viewpoint trajectory, Atlas is meant to generate corresponding views through a structured environment. For Li, that ability to direct the camera is what turns a generated scene into something users can shape and work within.

World Labs has described Atlas in a technical blog and is working to release it as a product, Li said. The company’s framing is broad rather than tied to one narrow application: spatial intelligence can apply to design and creative work, robotics, healthcare and education.

The robotics case is a shortage of usable environments

Fei-Fei Li says World Labs is a spatial-intelligence company, and argues that world models are useful wherever systems need to understand or construct physical space. The company is engaging users in creative and design industries as well as robotics-related fields, including healthcare and pharmaceutical laboratories. In each setting, Li emphasizes human control: creators can use the model as an augmentative tool, while developers can build downstream technology around their own requirements.

For robotics, the practical problem is building environments in which tasks and policies can be trained and evaluated. Li says that remains a bottleneck, particularly when robots must operate in real, human-centered settings. A robust world model could help by constructing worlds that can serve as training and evaluation grounds.

Pharmaceutical manufacturing and laboratory science are her concrete example. Such work is highly specialized, she says, and robots do not have much data from which to learn. World models could simulate those settings and turn the simulations into realistic places to develop and evaluate robotic systems.

World Labs illustrated the idea with footage of a robot arm placing a small green ball onto a wooden marble-run track, followed by a detailed simulated rendering of the same task and environment. Another demonstration showed a robot arm handling boxes, alongside reconstructed environments including an outdoor patio, a living room and a construction site. The demonstrations positioned simulation not as a visual effect but as infrastructure for training and testing physical systems.

Evaluation cannot stop with the companies building the models

Fei-Fei Li says World Labs runs internal technical benchmarks that consider both model capability and safety. But she argues that evaluation and benchmarking for increasingly capable AI must be an ecosystem-level responsibility, rather than something companies can resolve through their own tests alone.

Her proposed participants are specific: independent bodies, academia, government, industry, the public sector and civil society. When Ludlow asked whether she envisioned a regulatory body or another format, Li answered, “all of the above.” She pointed to her role co-founding Stanford’s Human-Centered AI Institute before ChatGPT’s release. The institute formed a Language Model Research Center to publish benchmarks; Li also cited ImageNet, introduced 15 years ago, as one of AI’s early benchmarks.

The governance position follows her broader view of AI risk. Li does not treat world models as a separate safety category from language models. The same basic calculus applies across technologies, she says: tools can be beneficial, but their development and deployment must put safety and responsibility first. She compares AI with electricity, steam engines and modern transportation—not to diminish AI’s stakes, but to argue that society must hold two aims together: human safety and technological advances that help people.

On existential risk, Li does not offer a probability estimate. She places the decisive question in human conduct: how people wield technological power, govern it collectively and use it to improve their lives. Her response to whether humans are in control is normative rather than predictive.

AI, it’s not about AI. It’s about humans. It’s about our shared as well as individual dignity, agency, as well as our responsibility.

Fei-Fei Li

Li rejects what she calls fatalism. Because AI is a “civilizational technology,” she says, responsibility cannot be left to a single company or an assumed trajectory of technical progress. Internal testing remains necessary, but it belongs within a broader arrangement of independent benchmarks and shared institutional participation.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free