Two things are happening in AI at the same time, and most people have only met one of them.
The one you have met is the chatbot. ChatGPT, Claude, Copilot, Gemini. A model that takes what has been said and predicts what is likely to come next. Do that at scale on most of the readable internet and you get systems that write, summarize, translate, plan, and reason across almost anything expressed in words, code, or images. That family is called large language models, and it is what nearly every AI headline of the last three years has been about.
The one you probably have not met is a different class of model that represents a bounded environment and estimates how its state may change, particularly when an action intervenes. A robot arm reaching for a coffee cup. A driving system facing an unusual cut-in. A factory line with an unexpected part in the wrong position. A short interactive scene rendered on the fly for someone learning a procedure. Those systems increasingly rely on what researchers call world models.
The chatbot is the AI that continues your sentence. A world model is the AI that tries to rehearse possible futures inside a scene.
That distinction is the whole subject of this four-part series. It is worth the next fifteen hundred words, because right now it is the reason otherwise sensible AI conversations at the executive level keep talking past each other.
The mental model, before the taxonomy
Here is the pattern I lean on for the rest of the series, and I want to label it before I use it.
Language models are trained to predict the next token — the next fragment of a word, roughly — given the sequence so far. Do that at scale and you get systems that continue and coordinate symbolic information across an enormous range of tasks, including tasks that involve reasoning about consequences in words.
World models are trained to predict how a scene changes, often conditioned on the action about to be taken inside it. Do that on video, sensor traces, simulation rollouts, or robot trajectories, and you get systems that can generate, forecast, or plan against short trajectories in an environment rather than in an inbox.
That is the mental model. It is not a clean architecture taxonomy, and I want to say that out loud because the picture in the primary literature is messier than the split above implies.
Modern language models carry rich implicit world knowledge and reason about consequences in language. Several systems being sold as world models — for example Meta's V-JEPA 2 — build their predictions in latent representation spaces that would not look out of place inside a language model. Industrial platforms such as NVIDIA's Cosmos 3 explicitly package physical reasoning, world generation, and action generation into one model family. The boundary is porous, and the useful thing is to keep the two prediction problems distinct without pretending they map onto two mutually exclusive product categories.
Keep the two purposes clear, and the current landscape stops feeling contradictory and starts looking like two frontier problems that are already composing.
World models
Two prediction problems lead to different kinds of systems
Language-centered models continue representations. World models estimate how a bounded environment may change, often under action.
Two parallel prediction problems are shown. The first continues context into a token and response. The second combines an environment state with an action to estimate a possible next state.
01
prediction targetContinue the representation
Given language, code, images, or instructions, estimate a useful continuation.
What fits next?
02
prediction targetPredict the transition
Given an environment state and a possible action, estimate how the state may change.
What may happen next?
What a language model does well
The reason chatbots have swept the last three years is not mysterious. Enterprise work is unusually language-heavy. Emails, memos, meeting notes, tickets, policies, contracts, code review, product copy, marketing briefs, financial commentary. Most of what a large organization produces is text about text.
A model that continues and coordinates symbolic information sits on top of that workload well. It reads what has been said, drafts what needs to be said next, and does it in the tone you specify. It reasons about constraints and policies expressed in language. It orchestrates tool calls and multi-step plans. It is why large language models went from research curiosity to line-item budget in eighteen months.
Where language models struggle is exactly the boundary you would expect. They are strong at what has been said. They are weaker at grounded predictions about how a specific physical situation will evolve under a specific action, when the answer depends on details never written down. Ask a chatbot to describe how a robot arm should pick up a mug and it will produce a very good description. Ask it to actually get the mug into the hand of a moving human on a factory floor, reliably, and the description does not close the gap on its own.
That is not a criticism. It is the shape of what the model is optimized to be good at, and it is why the operator-facing end of the frontier is investing in a second family of models.
What a world model is being trained to do
The lineage is older than the current news cycle. Ha and Schmidhuber's World Models paper in 2018 formalized the vocabulary most researchers still use, and the Dreamer line has been improving the practical version through several generations. The most recent step, the Dreamer 4 project in September 2025, trained an agent in a learned Minecraft environment using offline data and reported real-time interactive inference on a single H100. That is a demonstrated result inside a game world. It does not, on its own, generalize to a real warehouse.
What is different this year is that the topic stopped being niche. Several groups shipped systems that publicly use the world-model label, each doing a materially different job.
Google DeepMind released Genie 3 in August 2025. DeepMind reports interactive 720p environments at 24 frames per second, with consistency over several minutes and visual memory reported up to roughly one minute. That is a vendor-reported capability, not evidence of industrial reliability.
Meta released V-JEPA 2 in June 2025. The paper reports pretraining on more than one million hours of video, with the V-JEPA 2-AC variant using less than 62 hours of robot video for zero-shot planning experiments. Success rates vary materially by task and object.
Wayve released GAIA-2 in March 2025 as a controllable, multiview driving-scenario generator, positioned by Wayve as an offboard development and evaluation tool rather than a proof of on-road autonomy.
World Labs released Marble in November 2025 and previewed Atlas in early access in September 2026 as tools that produce persistent 3D representations from text, images, video, and coarse 3D inputs. Runway released GWM Worlds 2 in September 2026, reporting indefinite interactive 720p/24fps video with audio and text-conditioned actions.
NVIDIA released Cosmos 3 in May 2026, described as one open model family spanning physical reasoning, world generation, and action generation.
Several groups. Several architectures. Reconstructive spatial models, generative video models, interactive generative environments, latent predictive models, and world-action stacks are doing structurally different jobs. All are being described in public materials as world models, which is exactly why the taxonomy matters more than the label.
Why an executive audience should care
If you run a business that lives in documents — banks, law firms, consultancies, most of professional services — much of your AI conversation for the next year is going to remain a language-model conversation. That is correct. That is where your data is.
If any material part of your business runs equipment, moves goods, operates a facility, drives a vehicle, controls a machine, or takes physical action in the world — logistics, manufacturing, energy, real estate operations, healthcare delivery, defense — then the world-model conversation is a second, distinct conversation worth having in parallel. It is early. It is not a rebrand. It is not the same conversation as the chatbot one.
That is the honest read, and I want to be careful not to hype it. The current systems produce impressive demonstrations. Independent evaluation of physical executability and downstream task utility is much thinner. Where it exists — for example RoboWM-Bench at CVPR Workshops 2026, WorldArena 2.0, RoboPhys-3D, and WoW-World-Eval — the picture is fragmented and task-specific. The direction is worth planning against. The pace and the shape of adoption are open questions.
Limitations, and the argument I would push back on
The strongest objection to this framing is that "world model" is a marketing label stretched across genuinely different systems, and that the label is doing more work in the discourse than the evidence supports.
That objection has real weight. The taxonomy emerging in the research literature — surveyed, for example, in World Action Models — already distinguishes latent predictive models, generative video models, interactive generative environments, reconstructive spatial models, and world-action stacks. Those categories do different jobs, and pretending they are one thing produces the kind of executive briefing where a demonstration of a 3D scene is presented as evidence that a robot will now be safer. The gap between visual plausibility and physical executability is the subject of Part 2, and I would not want anyone finishing Part 1 with the impression that the gap is small.
The version of this argument I would push back on, harder, is the reverse claim that world models are just a rebranding exercise and nothing structurally new is happening. The V-JEPA 2 paper, the Genie 3 reports, and the Dreamer 4 project describe systems that are architecturally distinct from language-first stacks and are aimed at prediction problems the language-first path was not built to solve. That is an inference about direction. It is not proof of industrial autonomy. It is enough, on my read, to warrant an executive conversation, and not enough to warrant treating any current vendor demonstration as a shipped operational capability.
What the rest of the series will do
Part 2 takes on rehearsal — what it means for a system to explore possible futures before acting, and where visual plausibility, controllability, physical executability, and downstream task utility line up or fail to line up.
Part 3 takes on substrate — the linked operational record of observations, conditions, actions, and outcomes that adapting and evaluating these systems in a specific setting tends to require, and why document retrieval, however good, is insufficient by itself for evolving state and dynamics.
Part 4 takes on convergence — how language, prediction, simulation, action, and evaluation are composing into product-level systems, and what that implies for governance, portfolio choices, and the questions executives should be asking before the loop shows up in the operating budget.
For now, the takeaway is small and specific. If your whole AI mental model is a chatbot, you are describing one of two prediction problems. The other one is being worked on in the open, by people with different training data and different hardware assumptions, and the vocabulary you have been using will need to stretch to include it.
Companion deck
The 24-slide World Models primer is available as PowerPoint and PDF. Claims are sourced as of 2026-09-10.
Sources
- Ha, David, and Jürgen Schmidhuber. "World Models." arXiv preprint, March 2018. Source.
- LeCun, Yann. "A Path Towards Autonomous Machine Intelligence." Open Review position paper, 2022. Source.
- OpenAI. "Video Generation Models as World Simulators." Technical report accompanying Sora, February 2024. Source. Vendor-reported framing of the world-simulator ambition.
- Google DeepMind. "Genie 3: A New Frontier for World Models." August 2025. Source. Vendor-reported interactive-generation capability.
- Meta AI. "V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning." Paper, June 2025. Source. Success rates are task- and object-specific.
- Danijar Hafner et al. "Dreamer 4." Project page and paper, September 2025. Source. Minecraft results should not be generalized to robotics.
- Wayve. "GAIA-2." March 2025. Source. Vendor-reported offboard scenario generator.
- World Labs. "Marble." November 2025. Source. World Labs. "Atlas." September 2026 early access. Source. Vendor-reported spatial generation and reconstruction.
- Runway. "Introducing GWM Worlds 2." September 2026. Source. Vendor-reported indefinite-interaction claim.
- NVIDIA. "Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3." May 2026. Source. Vendor-reported converged model family.
- Jiang et al. "RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation." CVPR Workshops 2026. Source.
- "WorldArena 2.0." arXiv preprint, May 2026. Source.
- "RoboPhys-3D." arXiv preprint, August 2026. Source.
- "WoW-World-Eval." arXiv preprint, January 2026. Source.
- Xu et al. "World Action Models: A Survey." arXiv preprint, May 2026. Source.
