Artificial intelligence is rapidly becoming better at understanding and generating text, images, audio and video. However, the next major step could be teaching AI to understand how the real world works, how objects move, how situations change over time and what could happen after a particular action.
This is where World Models come in.
A World Model is an AI system designed to build an internal representation of an environment and use it to predict or simulate possible future situations. The technology is becoming particularly important in areas such as robotics, autonomous vehicles, Physical AI, simulation, gaming, virtual reality and scientific research.
Companies and research organisations including Google DeepMind and NVIDIA are developing technologies related to World Models. Google DeepMind’s Genie 3 focuses on generating interactive environments, while NVIDIA’s Cosmos platform is designed around World Foundation Models for Physical AI.
World Models are still an active research area, and today’s systems cannot create a perfect digital copy of the real world. However, the technology could become an important bridge between today’s digital AI and future physical AI systems.
What Is a World Model?
In simple terms, a World Model is an AI system that tries to learn how an environment works and what may happen when something changes inside that environment.
For example, suppose an AI system sees a road.
A conventional computer-vision system might identify:
“There is a road, a car and a pedestrian.”
A World Model aims to go further.
It could attempt to understand:
“The car is moving forward. The pedestrian may cross the road. If the pedestrian moves into the vehicle’s path, the vehicle may need to slow down or stop.”
So, the basic idea is:
Traditional AI: What is happening?
World Model: What is happening, how is the environment changing, and what could happen next?
This distinction is not absolute because modern AI systems can already perform some forms of prediction and simulation. “World Model” is also a broad research term used for different architectures and applications.
How Does a World Model Work?
A World Model can be understood through four basic stages.
1. AI Observes the Environment
The system can receive information from different sources, including:
- Video
- Images
- Text
- 3D data
- Camera feeds
- Sensors
- Robot actions
- Driving data
- Simulated environments
For Physical AI, the goal is to provide AI systems with information that represents how the physical environment looks and changes.
NVIDIA’s Cosmos platform, for example, is designed around World Foundation Models that can work with physical-world data for applications such as robotics and autonomous machines.
2. AI Builds an Internal Representation
The system then attempts to understand the environment.
Imagine that a robot is placed inside a room.
The AI may need to understand:
- Where the table is
- Where the chairs are
- Where people are standing
- Which objects can move
- Which objects are obstacles
- How objects are positioned relative to each other
- How the environment changes over time
This is more than simply recognising objects in an image.
The system is trying to understand the relationships and behaviour of the environment.
3. The AI Predicts Possible Future Situations
This is one of the most important ideas behind World Models.
Suppose a robot is asked to pick up a bottle from a table.
The system could model a sequence such as:
Robot approaches table → identifies bottle → moves its arm → reaches the bottle → grips it → lifts it.
If an obstacle appears in the robot’s path, the system may need to consider an alternative route.
In this way, a World Model can potentially act like a virtual environment for testing possible actions before they are performed in the real world.
4. The AI Selects or Supports an Action
Once different possible outcomes have been modelled, an AI agent or robot can use that information to decide what action to take.
This creates a potential loop:
Observe → Understand → Predict → Act → Observe Again
This type of loop is particularly important for autonomous machines.
A Simple Example: Self-Driving Cars
Consider an autonomous vehicle approaching an intersection.
The vehicle detects:
- A car ahead
- A cyclist on the left
- A pedestrian on the right
- A changing traffic signal
A basic perception system can detect these objects.
A World Model-based system could potentially go further by modelling different future situations.
Scenario 1
The pedestrian remains on the sidewalk.
Possible result: The vehicle continues.
Scenario 2
The pedestrian starts crossing.
Possible result: The vehicle slows down.
Scenario 3
The pedestrian suddenly enters the vehicle’s path.
Possible result: The vehicle needs to brake or change its trajectory.
The goal is therefore not simply to recognise the current scene but to model how the scene could evolve.
This is one reason World Models are attracting attention in autonomous driving research.
Google DeepMind’s Genie 3
One of the most prominent examples in the World Model field is Google DeepMind’s Genie 3.
Google describes Genie 3 as a general-purpose world model capable of generating photorealistic environments from text descriptions and allowing those environments to be explored interactively.
Google introduced Genie 3 in August 2025.
For example, instead of asking an AI to create a single image of:
“A vehicle travelling through a mountain road during rain.”
A World Model can aim to generate an interactive environment containing the road, surroundings and changing conditions.
The important difference is that the environment is designed to be interactive rather than simply being a static picture.
Project Genie
Google DeepMind also introduced Project Genie, an experimental prototype built around Genie 3.
The concept allows users to create, explore and remix interactive worlds using text or images.
This points toward a possible future where AI-generated environments are not just images or videos.
Instead, users could potentially create:
AI-generated interactive worlds.
For example:
“Create a futuristic city with roads, buildings, vehicles and pedestrians.”
The resulting environment could then be explored or used for experimentation.
What Is NVIDIA Cosmos?
Another major development is NVIDIA Cosmos.
NVIDIA describes Cosmos as a World Foundation Model platform designed for Physical AI.
Its target applications include:
- Robotics
- Autonomous vehicles
- Industrial machines
- Physical AI systems
One of the important ideas behind Cosmos is that developers can use generated or simulated physical-world data to help train and develop AI systems.
This could reduce the need to collect every possible situation directly from the physical world.
Why Are World Models Important for Robotics?
Training a robot in the real world can be expensive and slow.
Imagine a company developing a warehouse robot.
The robot may have to learn how to operate around:
- Shelves
- Boxes
- Workers
- Forklifts
- Moving objects
- Different lighting conditions
- Unexpected obstacles
Testing thousands of situations in a real warehouse could require significant time, equipment and money.
A World Model and simulation environment could potentially generate many of these scenarios virtually.
The robot could then be trained or evaluated in simulated environments before being deployed in the physical world.
This is one of the major reasons World Models are closely connected with Physical AI and robotics research.
World Models and Humanoid Robots
Humanoid robots are another major potential application.
A future household robot might receive an instruction such as:
“Take the cup from the table and put it in the kitchen.”
To complete the task, the robot may need to:
- Find the cup.
- Understand the room.
- Plan a route.
- Avoid obstacles.
- Reach the cup.
- Grab it.
- Carry it safely.
- Find the kitchen.
- Put the cup down.
A World Model could potentially help the robot understand the environment and predict the consequences of different actions.
This is one reason World Models could become an important component of future general-purpose robots.
World Models and Autonomous Vehicles
World Models are also relevant to the development of autonomous vehicles.
Future autonomous vehicles may need to understand much more than:
“There is a vehicle ahead.”
They need to understand how the entire road environment is changing.
For example:
- A pedestrian may start crossing.
- A cyclist may change direction.
- Another vehicle may suddenly brake.
- A traffic light may change.
- Road construction may alter the normal driving path.
A World Model could potentially help an autonomous-driving system simulate possible outcomes and improve decision-making.
NVIDIA’s Cosmos platform specifically targets physical AI applications including autonomous machines.
World Models Are Not Limited to Robotics
World Models could potentially have applications across many industries.
Autonomous Driving: Simulation and prediction of complex traffic situations.
Robotics: Training robots and testing actions in virtual environments.
Manufacturing: Simulating factories, machines and production processes.
Gaming: Creating dynamic and interactive game worlds.
Virtual Reality: Generating environments that users can explore.
Healthcare: Potentially modelling complex environments and biological processes.
Smart Cities: Simulating traffic, infrastructure and energy systems.
Space: Simulating planetary environments, spacecraft operations and mission scenarios.
Scientific Research: Modelling complex physical systems and generating experimental scenarios.
World Model vs Digital Twin
World Models and Digital Twins are related but they are not exactly the same.
Digital Twin
A Digital Twin generally represents a specific real-world object, machine, facility or system in digital form.
For example:
A factory → digital representation of that factory
World Model
A World Model focuses more broadly on representing an environment and modelling how that environment can change.
For example:
Factory + robots + workers + machines + objects + interactions
A simple way to remember the difference is:
Digital Twin = Digital representation of a specific real system
World Model = A model of how an environment and its elements can behave or evolve
In practical applications, both technologies can be used together.
Why Do World Models Need So Much Data?
Understanding the physical world is much more difficult than understanding text.
A World Model may need to learn from:
- Videos
- Images
- 3D environments
- Sensor data
- Robot movements
- Driving footage
- Physical interactions
- Simulated environments
For example, if an AI needs to understand how an object falls, it needs information about movement, gravity, surfaces and interactions.
This makes World Model training computationally expensive.
NVIDIA’s research around Cosmos focuses heavily on large-scale physical-world data and infrastructure for developing Physical AI systems.
The Biggest Question: Does AI Really Understand Physics?
One of the biggest research questions around World Models is whether an AI system actually learns the underlying rules of the physical world or simply learns statistical patterns from its training data.
Consider a simple example.
An AI sees thousands of videos of balls falling.
It may learn:
“When the ball is above the ground, the next frames usually show it moving downward.”
But does it actually understand gravity and physical causality?
That distinction becomes important when the AI encounters a situation it has never seen before.
Researchers therefore study issues such as:
- Physical consistency
- Causal reasoning
- Spatial reasoning
- Generalisation
- Object permanence
- Long-term prediction
- Real-world reliability
For physical AI, simply generating a visually realistic environment is not enough. The environment also needs to behave in a useful and consistent way.
Another Major Challenge: Long-Term Prediction
Predicting what will happen a few seconds from now can be easier than predicting what will happen much further into the future.
For example:
Current situation → 1 second later
may be relatively manageable.
But:
Current situation → 10 minutes later
can become much harder.
Small errors can accumulate.
If an AI predicts that a vehicle is slightly to the left of its real position, that error can influence the next prediction. After many steps, the simulated environment could become significantly different from reality.
Therefore, long-term consistency remains an important research challenge.
Can a World Model Create a Perfect Copy of the Real World?
Not today.
Current World Model systems are impressive, but they are not perfect digital copies of reality.
There are still challenges involving:
- Physical accuracy
- Long-term consistency
- Real-time generation
- Control
- Compute requirements
- Unexpected situations
- Real-world transfer
Google has also described limitations around Genie 3 and its current capabilities.
Therefore, World Models should currently be considered an active research and development area, rather than a finished technology.
World Models and Generative AI
Generative AI today is primarily associated with creating:
- Text
- Images
- Audio
- Video
World Models take the concept in another direction.
A simplified progression could be:
Generative AI → Generate Content
World Model → Generate and Simulate Environments
AI Agent + World Model → Plan and Act in Environments
Physical AI + World Model → Act in the Real World
This is a conceptual technology relationship rather than a fixed industry roadmap.
World Models and AI Agents
AI agents are designed to perform tasks rather than simply answer questions.
An agent operating in a complex environment needs to understand:
What is happening now?
But it may also need to answer:
What could happen if I take this action?
A World Model could potentially provide that predictive layer.
For example, imagine an AI agent operating inside a warehouse.
It needs to deliver a package.
It could potentially consider:
Route A: Short but blocked.
Route B: Longer but clear.
Route C: Fastest under current conditions.
The agent could then select an action based on its objectives and the model’s predictions.
This is one potential way World Models and AI agents could work together.
Who Is Working on World Models?
Google DeepMind
Google DeepMind is one of the most visible organisations working on World Models.
Its Genie 3 focuses on generating interactive environments, while Project Genie demonstrates an experimental interface for creating and exploring such worlds.
NVIDIA
NVIDIA is developing Cosmos, a platform focused on World Foundation Models for Physical AI.
Its target areas include robotics and autonomous machines.
World Labs
World Labs is working on AI systems related to 3D world generation and spatial intelligence.
Universities and Research Labs
Universities and research institutions around the world are also researching:
- World modelling
- Embodied AI
- Simulation
- Spatial intelligence
- Physical reasoning
- Robotics
The field is therefore not controlled by a single company or country.
What Could World Models Look Like in the Future?
A simplified progression could look like this:
Today’s AI
“This is a car.”
More Advanced AI
“The car is moving toward the intersection.”
World Model
“If the car continues moving at this speed, it may reach the intersection in a few seconds.”
More Advanced Physical AI
“Here are several possible outcomes based on the current environment.”
Future Autonomous System
Observe → Predict → Plan → Act → Learn
This is why World Models are being investigated as a potential foundation for future Physical AI systems.
Potential Impact of World Models
| Sector | Potential Impact |
| Autonomous Vehicles | Simulation of complex driving scenarios |
| Robotics | Virtual training and environment understanding |
| Humanoid Robots | General-purpose physical tasks |
| Manufacturing | Factory and machine simulation |
| Gaming | Dynamic AI-generated worlds |
| VR/AR | Interactive digital environments |
| Healthcare | Medical and biological simulation |
| Space | Mission and planetary-environment simulation |
| Science | Modelling complex physical systems |
| Smart Cities | Traffic and infrastructure simulation |
Many of these applications are still under research or development, so they should not be treated as guaranteed future products.
Why Are World Models Important?
The biggest challenge for the next generation of AI may not simply be producing better text or images.
It may be teaching AI to understand the physical world well enough to interact with it reliably.
World Models are one approach toward solving this problem.
The basic concept can be summarised as:
Observe → Understand → Simulate → Predict → Act
Google DeepMind’s Genie 3 demonstrates the direction of interactive AI-generated environments, while NVIDIA’s Cosmos demonstrates the use of World Foundation Models for Physical AI development.
The technology still faces significant challenges, particularly around physical accuracy, long-term consistency, computational cost and reliable transfer from simulation to the real world.
However, if these challenges can be addressed, World Models could become an important technology connecting today’s Generative AI and AI Agents with tomorrow’s robots, autonomous vehicles and Physical AI.
Outcome
World Models represent a shift from AI that mainly processes information to AI that attempts to understand how an environment behaves and what could happen next.
Instead of simply recognising a car, an AI system could eventually model where that car might move. Instead of simply identifying a cup, a robot could potentially understand how to reach it, pick it up and move it without hitting surrounding objects.
The technology is still evolving, but research from companies such as Google DeepMind, NVIDIA and academic institutions shows that AI-generated environments, simulation, spatial intelligence and Physical AI are becoming increasingly connected.
For the future of robotics, autonomous vehicles, gaming, simulation and intelligent machines, World Models could become one of the important building blocks of the next generation of AI.






























































