105 Comments
User's avatar
Bryan Caballero @ The Shield's avatar

The most interesting idea is that world models may eventually force AI research into the same problem humans and institutions already struggle with:

not prediction, but referential integrity.

Because once a system begins operating through internal representations of reality, a deeper question emerges:

> How does the model know its world still corresponds to the originating conditions that made the model valid in the first place?

That feels enormously important.

A model that never updates drifts toward delusion.

But a model that updates continuously without continuity risks fragmentation and loss of identity.

So the challenge may not simply be: “Can a system build a world model?”

But:

> “Can it preserve meaningful alignment between representation and reality as recursion, scale, speed, and self-modification increase?”

Which may ultimately make:

friction,

embodiment,

memory,

temporal continuity,

and relational interaction

far more important than pure optimization.

Because eventually the danger is not merely incorrect prediction.

It’s when the proxy quietly becomes more operationally influential than the originating reality itself.

Paul Lorsbach's avatar

Agreed, and maybe what's missing (or what's worth calling out more explicitly) is the full feedback loop back into the world. Planning has to result in action in the real world, which is then correlated back through the render->simulate->plan stack. If the agent doesn't recognize that a planned action had an interpretable consequence, the best a plan can do is select a most probable future without itself in it, like watching a movie. That's just prediction.

That doesn't solve the problem, it's just necessary. You can still drift, and that's probably an intractable problem. If all 3 layers of the stack are updatable based drift from what was expected, then you have the best chance at staying cogent with the world, but you by definition can't have a perfect model.

Dean of Big Data 🎓 #DOBD's avatar

Do you really mean:

Where language models learn the statistical structure of text, world models learn the *physics* of space and time?

Using the word "statistical" in this statement - "Where language models learn the statistical structure of text, world models learn the statistical structure of space and time" - doesn't make any sense.

Alex Inch's avatar

In the POMDP formalism that Fei-Fei references, it's fair to call space and time stochastic from the perspective of a limited agent.

When you turn your head, you can't causally derive what's happening behind you — you can only make a good guess by extrapolation. There's no privileged observer that perceives the "true" state of reality. Agents get a limited scan of information and have to construct their best guess of what the wider world state is. When that prediction is wrong, we exhibit surprise.

Birchfabric's avatar

This is true, unless you move your hand behind you head. You do not have to see your hand to know it's there, which is quite interesting to consider and gives a hint at what's required for AI, in my opinion.

Adam Murray's avatar

Alex I agree that an output we experience in life is “surprise”. A Bayesian feedback loop a prediction error. However I think the more important surprise is not the prediction error but that we just don’t know or may not know. A better model will be made to fix a prediction error surprise.

“what is”, in Li’s clean phrase, rather than what the viewer happens to see, about a being who by the very same picture can never stand at that vantage toward herself.

And part of what is true about a person, the part that does not survive the trip into state, is that her own future is open to her.

The not-yet-knowing is not a gap in the data. It is a property of the thing. Close it and you have not completed the map. You have changed what was on the table from someone into something. Again we are not a cup.

Paul Lorsbach's avatar

I think it's a fair statement. The "physics of space and time" is just statistical model, one governed by quantum mechanics. Arguably, even though the stochastic nature of physics at the levels we're working are rounding errors, they are fundamentally the reason why we can't actually predict and plan the future.

Dean of Big Data 🎓 #DOBD's avatar

Paul, not sure I agree. The physics of space and time are driven by causal factors, where you can quantify cause-and-effect and from which you can conduct interventions and "what if" scenarios to see the impact of changing the causal factors on the outcomes. They are not statistical factors that rely on correlation-based relationships.

I mean, that is the foundation behind the entire digital twins concept.

Paul Lorsbach's avatar

Yeah I'm in agreement with you on that point, I think we're probably just quibbling about the scope and scale at which we can count on those causal factors to be absolutely true. A camera capturing a scene can only do so at a certain fidelity, it has to drop some information that otherwise existed or was observable at the moment. The simulations can not have absolute understanding of the physics of the model. Even if it's encoded a formula for proven physics, there are still precision errors with calculations. Plans are the same. If the full system is built well enough, and the noise bounded well enough, you can get pretty dang close to a perfect approximation. But of course, it's can't be perfect, and even getting close is often the whole challenge.

Chris Rowe | Signal & Noise's avatar

I respect the ambition and clarity of the work you’re presenting here, and I say this as someone who has spent a lot of time at the interface of technical work, government, and communication.

My concern is not the research itself but the way it is being surfaced. Spatial intelligence and rich world‑models feel less like “interesting rabbit holes” and more like black holes: once you really follow them, they pull in everything—war, infrastructure, bodies, stories—not just software. That is precisely why I’m uneasy seeing this framed as broad thought‑leadership content on Substack and LinkedIn.

In my view, before we talk about “integrating” humans with increasingly capable AI, we need humans who are genuinely integrated with themselves. That dimension is largely absent from the public discourse. Many readers here will understandably see “wow, singularity!” without much grounding in what it would mean to be ready for that as a person, a community, or a country.

I’m not arguing against the work. I’m suggesting that some of these ideas may belong first in deeper, more bounded rooms—academic, policy, and serious leadership contexts—before they are packaged for general feeds. Otherwise we risk mistaking virality for understanding and skipping the human‑integration work that should come first.

Respectfully,

Chris Rowe

Signal & Noise

JQ's avatar

Isn't 'academic' often more "supposed" to be free and democratized than not, unless of course if you're talking about stuff like uranium bomb making, but this at this stage is not comparable to that, unless you think it could be.

"Serious leadership" not always and necessarily they can be trusted to this.

On one hand you want dissemination of information and knowledge to spark the next cohort of bright researchers to emerge and participate and all that, on the other hand you may risk bringing in or leaking to bad actors. I'd suggest to people who are genuinely more security- and risk-centri to dive fairly deep in research - understanding the actual maths and operations and capabilities and limitations - rather than holding a strong limitation-first stanc, trying to limit AI research and more importantly trying to limit dissemination of information on the latest ideas and research on AI.

I do not disagreeing to what you address, however I do largely disagree that "these ideas belong first in deeper, more bounded rooms". Yeah and I get that you meant "SOME of these ideas MAY".

"we need humans who are genuinely integrated with themselves"

Always true. However, who is to say who is integrated. I think this responsibility expands to the wider society rather than falls onto the shoulders of the existing researchers. If we have more integrated people in the society, then chances are we have more integrated people in AI research, in fact we will have more integrated people in pretty much any field in a highly positively correlated manner.

Chris Rowe | Signal & Noise's avatar

Thanks for taking the time to write such a thoughtful response.

I’m very much in favor of broad, democratized access to academic work and the kind of open circulation that brings the next cohort of bright people into the field. That’s how I got here too, and I don’t want to see that impulse die.

Where I’m probably more cautious is around context and audience, especially for work that is unusually “operationalizable” or dual‑use. I’m not arguing for a blanket secrecy regime or saying “serious leadership” is automatically trustworthy; I’m saying that some ideas initially benefit from being worked through in rooms where you’ve got the right mix of technical depth, risk literacy, and responsibility on hand before they’re packaged for a wider, more decontextualized audience.

On the “integrated humans” point, I agree this isn’t something a narrow priesthood of researchers can or should arbitrate alone. You’re right that it’s ultimately a broader societal project: the more integrated people we have in general, the more of them will end up in AI, in policy, in business, in everything else. My concern is just that, in the meantime, the people closest to these systems don’t shrug off their own particular responsibilities because “society” will catch up later.

So I don’t think we’re as far apart as it might sound. I’m for open inquiry; I’m also for being honest that some ideas carry higher stakes than others, and that where and how they’re first shared can matter a lot.

Mariangela Salafia's avatar

What struck me reading this is that renderers, simulators and planners all operate within the agent-world loop.

Yet in most organisational settings, another layer sits on top of that loop: deciding which actions are authorised, accountable and worth taking.

As our ability to model the world improves, the limiting factor may increasingly become our ability to make and govern decisions within it.

World models help us understand reality. Organisations still need mechanisms to convert that understanding into accountable action.

Paul Lorsbach's avatar

This is it! This is the crux! If the model cannot also model how it's plan actually hits the world, it'll be forever knee capped. At some point, we can't be the makers of decisions, we have to let the models, we need to act as gardeners or parents. They stop being one directional lens or tools for us, and start owning their actions, and we'll learn more from their outcomes than we will from their inference.

nihal | deeptech decoded's avatar

The world is not made of words. – Loved that first line to set the tone. Thankyou!

Harald Schepers's avatar

Words are the abstract representation of the World for the human kind

nihal | deeptech decoded's avatar

Indeed but also quite an efficient and effective way of communicating with one another. :)

Harald Schepers's avatar

communication is the way how language will develop, how knowledge will improve and grow - at least presently.

Mañana's avatar

Very helpful overview. And good to see Wittgenstein mentioned albeit as a foil. He also said the world is the totality of facts not things.

Joe Micallef's avatar

I'm writing this comment as I wrap up teaching an animation class at Pasadena City College. Just the week before, I finished your book, *The Worlds I See*. I was very interested in your journey through computer vision and AI, but ultimately, what interested me most was the connection between vision and imagination. In an artist's imagination, a world can be anything, and because I have a background in fine art, CG animation, and advanced manufacturing, I see your World Model Taxonomy as an opportunity to bridge art and engineering.

In a unified world model, I envision simulations as worlds, making such a model ideal for studying complex adaptive systems and opening up the imagination to the idea that a "world" can be anything: an atom, a solar system, the circulatory system, a bacterial ecosystem, or something entirely made up and abstract.

I'm very interested in using ML, USD, and procedural simulation tools such as Houdini to create adaptable simulations that exist in splat-based environments. I see Houdini, along with many other CG tools for students, combined with Marble from World Labs, as a valuable entry point because it offers a process that anyone can learn as a first step toward developing datasets for world models. Having students and researchers build those processes could be very rewarding. This is where we bridge the gap between art and science, using world-model simulations to research complex problems and drive real societal improvements.

As a teacher, I'm becoming involved in an AI consortium within the community college system here in Southern California. Much of the conversation focuses on LLMs, and among the 200 educators in the consortium, I believe I may be the only one focused on world models. Thank you for presenting the taxonomy. I will share it with my fellow educators because it provides a crucial understanding that the future of AI extends beyond LLMs, and that students in the arts, engineering, education, and science will play a crucial role in developing worlds that suit their needs and areas of expertise.

Kevin McLeod's avatar

The brain does not render, uses no models, it does not use symbols or words to think.

This is simply more wasted compute that has no relationship to brains or reality. It’s another pointless rabbit hole.

Dorian's avatar

world is not made of words.

But it is not made of models either.

Every model is a compression algorithm.

The real question is not how accurately it predicts.

It is what gets discarded during compression.

Jiada Li's avatar

One of the clearest clarifications for Renderers, Simulators, Planners, and the Loop That Connects Them that I've ever seen!

Nick Kaufmann's avatar

Reality always overcomes the plan

Georgi Paleshnikov's avatar

People plan, God decides... an old Bulgarian proverb...

Robert Koller, CAIA's avatar

The loop defines state as physical reality, the geometry, the velocities, the forces, and the planner derives its actions from that. There is a second state a planner in any consequential setting also has to track, and the loop is silent on it: what the agent is permitted to do, not only what is physically possible.

A simulator tells you the cup falls when pushed. It does not tell you whether you were allowed to push it. In a living room that gap is nothing. In an operating room, a cockpit, or a settlement system it is the entire problem. And unlike physics, this layer is not learned from observation. It is authored, in procedures, regulations, and law, which means it has to be compiled rather than inferred.

A planner that models only the physical substrate will produce actions that are physically valid and not permitted. Next to the simulator there has to be a model of the allowed, or the planner is reliable in the domains where reliability does not matter and unproven in the ones where it does.

The Bedrock Project's avatar

So strong on world representation, weak on governed execution.

perception

state

action

but not:

authority

policy

attribution

receipts

economic consequence

James Andrews's avatar

"Language models have given machines an extraordinary command of concepts, vocabulary, and reasoning..." This statement isn't precise.

Mañana's avatar

Imprecision is a bridge between the continuous world and discrete sentences. Think for a moment - if our words had rigid, mathematically precise borders, language would break the moment reality shifted even slightly. Imprecision gives language the flexibility to absorb new concepts and grey areas without needing an entirely new vocabulary.

James Andrews's avatar

Nice response ChatGPT... or is it Claude, I can't tell.

Mañana's avatar

Neither actually, but you don't need to explain that you can't tell - your comment has the hallmark of a human stochastic parrot!

JC's avatar

It's because it was written by an LLM..

Paul Topping's avatar

Perhaps she's just throwing a bone to the LLM makers in order to get them to keep reading.

Mañana's avatar

Precise enough ... anyway language isn't precise

James Andrews's avatar

Is that your "truth?"

Mañana's avatar

it's a feature of language

Hardik Kabaria's avatar

Render, simulate, and plan is the right structure. But world models need measurable physics to matter in the physical world.

That is the argument we laid out in our foundation model for physics (https://github.com/Vinci4d/continuous-physics-reasoning): physical-world AI needs models that can render, simulate, and plan (we call it optimize) real systems — not just represent them. We published it because we saw the same gap from the engineering side: world models cannot stop at representation; they need measurable physics.

At Vinci, we started by getting the simulator layer right because that is the qualification bar for engineering. If a physical-world model cannot compute measurable behavior with deterministic, solver-accurate fidelity, it cannot be trusted for real design decisions. We are running this on production semiconductor and hardware workloads today.

We are at a critical point in the world models conversation: rendering and planning are only as useful as the underlying physics layer. In engineering, approximate physics is not a limitation. It is a failure mode. If the model cannot resolve how matter, geometry, heat, stress, and constraints behave at the scale where real systems fail or succeed, it cannot be trusted to reason or act in the physical world.

We should put these two papers side by side and walk through what changes when the simulator layer is already working in production. Vinci is proving that layer today, and what we are seeing has direct implications for how physical-world models should be defined.

Think AI's avatar

Professor Fei Fei’s piece of information encourages faith in AI.