Fei-Fei Li and Yann LeCun Launch Billion-Dollar Ventures to Build 'World Models' for AI
Two of the most prominent figures in artificial intelligence have raised over $1 billion each to build 'world models'—AI systems designed to understand physical reality rather than just language.
By Sofia Matos
- Generative Spatial Advocates
- Believe world models must render high-fidelity 3D environments to achieve spatial intelligence.
- Abstract Reasoning Proponents
- Argue that predicting in a latent, mathematical space is the most efficient path to physical AI.
- AI Industry Analysts
- Focus on the massive capital shift away from LLMs and toward physical simulation.
The common assumption is that because artificial intelligence can write flawless code, pass medical exams, and generate photorealistic art, it understands the world. It does not. Large language models (LLMs) are fundamentally blind to physical reality. They do not know that a dropped glass shatters, that water flows downhill, or that two objects cannot occupy the same space. They only know the statistical probability of the words used to describe those events. This gap between linguistic mimicry and physical reality is why the current generation of AI struggles to operate robots or navigate complex environments without hallucinating physics. To move beyond text, the AI industry is undergoing a massive architectural shift toward 'world models'—internal simulators that predict how physical environments evolve over time.[1][2]
In early 2026, this shift accelerated when two of the most decorated figures in artificial intelligence each raised roughly $1 billion to build these physical simulators. Fei-Fei Li, the Stanford professor who pioneered modern computer vision through ImageNet, secured a $1 billion funding round for her startup World Labs, bringing its total capital to $1.23 billion and its valuation to an estimated $5 billion. Weeks later, Turing Award winner and Meta Chief AI Scientist Yann LeCun raised $1.03 billion for his new venture, Advanced Machine Intelligence (AMI) Labs, valuing it at $3.5 billion. Both founders are betting that the next frontier of artificial intelligence will not be won by feeding more text into language models, but by teaching machines the fundamental laws of cause and effect.[3][4][5][6][7]
While both founders use the term 'world model,' they are building fundamentally different mechanisms to solve the same problem. The core mechanism of any world model is predictive simulation: forecasting future states based on current actions and observations. If an autonomous agent decides to step on a wet floor, the world model must predict that the foot will slide, allowing the agent to adjust its plan before acting in the real world. This requires an AI to internalize concepts like gravity, inertia, friction, and object permanence. However, the way the AI represents this information divides the field into two distinct schools of thought: the generative school, which renders explorable 3D worlds, and the abstract school, which predicts outcomes in a mathematical space without generating pixels.[2][3]
Fei-Fei Li's approach at World Labs focuses on 'spatial intelligence' and the generative school of world modeling. Her systems are designed to render explorable, three-dimensional environments that humans and AI agents can interact with. World Labs recently released Marble, a model that generates interactive 3D replicas of spaces from visual or written prompts. By building AI that can perceive and generate 3D geometry, Li aims to create simulated environments where physical laws apply. These highly realistic virtual spaces can be used for architectural design, gaming, and crucially, training robots in simulation before they touch the real world. The strength of this approach is its visual fidelity, making it immediately useful for human creators and spatial computing applications.[1][3][6]
Fei-Fei Li's approach at World Labs focuses on 'spatial intelligence' and the generative school of world modeling.
Yann LeCun's AMI Labs takes a radically different approach, rooted in the Joint Embedding Predictive Architecture (JEPA). Instead of generating pixels or 3D visuals, LeCun's world models predict outcomes in an abstract, mathematical 'latent space.' LeCun argues that reality is continuous, noisy, and too high-dimensional to be solved by predicting every pixel of a video frame. Instead, a JEPA model strips away irrelevant visual details—like the exact texture of a falling leaf or the shifting shadows in a room—and focuses purely on the underlying physical dynamics and causality. This makes the model highly efficient for autonomous agents and robots that need to plan multi-step actions in real time, as they do not waste computational power rendering visuals they do not need.[2][3][4][5]
The stakes for this architectural divergence are immense. Whoever successfully builds the foundational world model will likely control the substrate upon which the next decade of physical AI is trained. If Li's generative approach wins, the future of AI will be deeply visual, merging simulation, media, and robotics into a single spatial computing paradigm where agents learn inside photorealistic digital twins. If LeCun's abstract approach proves superior, autonomous systems will operate on highly efficient, invisible representations of reality, prioritizing reasoning and planning over visual fidelity. Both approaches require massive capital, which explains why hardware giants like Nvidia and AMD, alongside major venture capital firms, are heavily backing both sides of the race.[3][4][6]
Despite the massive influx of capital, significant uncertainties remain in the quest to build a true world model. These systems are notoriously difficult to scale and stabilize. While they can generate structurally coherent spaces from limited data, they currently struggle to maintain logical object permanence and precise spatial reasoning over long time horizons. A model might perfectly simulate a room but 'forget' that a wall exists behind a closed door after a few minutes of interaction, leading to physical hallucinations. Furthermore, the transition from theoretical research to commercial application for these physical simulators could take years, requiring vast amounts of compute and entirely new training methodologies that move beyond the static datasets used to train language models.[1][2][5]
Ultimately, the emergence of World Labs and AMI Labs signals the end of the era where language models were viewed as the sole path to artificial general intelligence. By shifting the training target from the world of words to the world of things, Li and LeCun are attempting to give machines the one thing they currently lack: common sense. If successful, these billion-dollar ventures will unlock a new class of autonomous systems capable of safely and reliably navigating the physical world. This would transform industries from manufacturing and healthcare to autonomous transport, as robots would finally be able to anticipate the consequences of their actions before they take them.[2][3][4][5]
Key points
- Artificial intelligence is undergoing a massive shift from language-based models to 'world models' that understand physical reality.
- AI pioneers Fei-Fei Li and Yann LeCun have each raised roughly $1 billion for startups dedicated to building these physical simulators.
- Li's World Labs focuses on 'spatial intelligence,' generating highly realistic, explorable 3D environments for design and robotics.
- LeCun's AMI Labs uses an abstract architecture called JEPA, which predicts physical outcomes mathematically without rendering pixels.
- Both ventures aim to solve the critical limitation of current AI: the inability to reliably reason, plan, and act in the physical world.
Key terms
- World Model
- An AI architecture that learns the dynamics of physical reality—such as gravity, friction, and object permanence—to predict future states of an environment.
- Spatial Intelligence
- The ability of an AI system to perceive, understand, and interact with three-dimensional space and geometry.
- Latent Space
- An abstract, mathematical representation of data where an AI model processes information, stripping away unnecessary visual details to focus on core concepts.
- JEPA (Joint Embedding Predictive Architecture)
- An AI architecture championed by Yann LeCun that predicts the outcomes of actions in an abstract mathematical space rather than generating visual pixels.
- Next-Token Prediction
- The mechanism used by large language models to generate text by statistically guessing the most likely next word in a sequence.
Sources
[1]ForbesGenerative Spatial AdvocatesA discernible shift is occurring within AI research, moving from generative models for language and images toward the development of world models
Read on Forbes →
[2]AI.ccAI Industry AnalystsWorld Models in 2026: Why Google, NVIDIA, LeCun & Fei-Fei Li Are Betting Billions on AI That Understands the Physical World
Read on AI.cc →
[3]MediumAI Industry AnalystsWhy Fei-Fei Li and Yann LeCun each raised a billion dollars to bet against language
Read on Medium →
[4]Weights & BiasesAbstract Reasoning ProponentsYann LeCun raises $1B for world model startup
Read on Weights & Biases →
[5]TMVAbstract Reasoning ProponentsYann LeCun's AMI Labs raises $1.03B to build world models
Read on TMV →
[6]Silicon RepublicGenerative Spatial AdvocatesFei-Fei Li's World Labs raises $1bn to advance spatial intelligence
Read on Silicon Republic →
[7]StartupHubGenerative Spatial AdvocatesFei-Fei Li's World Labs: $1.23B Raised, Marble Now Shipping
Read on StartupHub →
Comments
Every angle. Every day.
Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.
