Earlier this month, World Labs, founded by AI pioneer Fei-Fei Li, released something called Atlas. An almost magical demo of that went viral: give Atlas a single photograph, then move a virtual camera around it, and you can look at the scene from the side, above it, behind it, even at parts the original camera never saw.
From one flat image, Atlas imagines a coherent three-dimensional world around it. Give it a few more images and imagination transforms to reconstruction. World Labs calls Atlas an “omni world model”, trained natively across text, images, video and 3D. Since ChatGPT burst onto our world four years ago, AI has been synonymous with LLMs, or large language models.
Every month since then has brought us more and more powerful LLMs, trained on trillions of parameters, costing billions of dollars. But just making LLMs more powerful is not the next frontier; the next inflection could be world models.
This is because the world is not a paragraph of words but exists in a glorious three-dimensional cornucopia of objects, sounds, sights and senses.
More than words
LLMs have acquired extraordinary capabilities by consuming enormous quantities of text, images and increasingly other media. But humans do not learn this way. A child does not read Wikipedia before understanding that a cup falls when pushed off a table, or that a ball rolling behind a sofa still exists. AI’s “pre-training data” is the textual and visual internet, but our pre-training dataset is reality itself.We see in three dimensions, we hear voices with tone and direction, we touch things, move through rooms, taste food and learn what happens when we act upon the physical world. Our developing brains continuously combine perception, movement, memory and prediction. Long before children acquire sophisticated language, they have absorbed an astonishing amount of common sense about objects, distance, gravity, other people, and cause and effect.
There is also a question of bandwidth. Typing communicates words, but speaking adds tone, rhythm, emphasis and sound, and that is why the best way to work with AI so far has been to speak with it. But vision adds infinitely more to it: objects, depth, movement and context, making the informational richness of an experience explode. Children do not need to be trained on the “entire internet”; they can learn remarkable things from comparatively little lived experience because each experience carries vastly more structure than a sentence.
Physical world
This is precisely the limitation that Yann LeCun, one of AI’s foundational researchers, has been attacking for years. He argues that human-level intelligence cannot emerge simply by feeding machines ever more human-written text. Current systems, he says, are extremely good at manipulating language, which can fool us into “equating eloquence with intelligence”. His alternative is AI that learns how the physical world works, predicts what will happen after an action and uses that prediction to plan. Li has staked her next chapter on a related idea: spatial intelligence.
Her argument is that while language intelligence lets machines understand and generate words, spatial intelligence lets them reason about objects, places, geometry, movement and interactions in the physical world. She calls this the transition “from words to worlds”.
So, what exactly is a world model? At its simplest, it is an internal representation of an environment that allows AI to understand its state, imagine how it might change and predict the consequences of actions. An LLM asks, loosely, “What token comes next?” A world model asks something closer to, “If this is the world now, and I do this, what will the world look like next?” This shift might sound small, but its complexity and implications are enormous.
Gaussian splat
A wonderfully named building block in this new vocabulary is the Gaussian splat. A token is a small unit in a stream of language. A Gaussian splat is a tiny, fuzzy, semitransparent blob positioned in 3D space.
Each carries information about where it sits, its shape, orientation, opacity and appearance. Millions of these splats can blend to form a photorealistic threedimensional scene. “Gaussian” comes from Carl Friedrich Gauss and the bellshaped mathematical distribution bearing his name.
“Splat” refers to how these fuzzy blobs are projected, or splatted, onto the screen to render the scene. Simplistically put, if tokens help machines talk about reality, splats help machines build and render it. Atlas can generate explicit 3D worlds as Gaussian splats, including geometry that no camera originally captured.
The future
Now imagine where this can go.
A robot needs a world model to understand a kitchen rather than merely recognise the word “kitchen”. Autonomous vehicles need to predict how pedestrians and traffic will move. Factories could create living digital twins and test changes before touching the production line.
Architects could walk through buildings before constructing them. Students could enter ancient Rome rather than read a paragraph about it. Personal AI could understand where things are in our homes, not merely what our calendars say. At global scale, the same conceptual direction matters for weather, climate, materials, biology and other systems where intelligence means modelling complex worlds evolving over time.
If language can describe a cyclone, a world model aspires to simulate what the cyclone might do.
The obstacles to building world models are formidable, as physical causality is much harder than visual realism. World models will need far richer data, better representations of physics and vastly more grounding in reality. But the direction is important.
This leap can make machines that can understand where they are, what surrounds them, what might happen next and what their actions will change. As Li puts it, “Language gave machines a way to talk about that world. World models are how machines will finally come to understand, imagine, reason and interact with it.”
In the beginning was the Word. AI learnt that first. Now it is reaching for the World.
(Views are personal)












Leave a Reply