Fei-Fei Li vs. Yann LeCun: The Growing Debate Over World Models
The road to AGI has finally converged on one battlefield—world models. Fei-Fei Li has unveiled Marbl...
The road to AGI has finally converged on one battlefield—world models.
Fei-Fei Li has unveiled Marble, the first commercial world model from her company.
Almost simultaneously, Yann LeCun announced his departure from Meta to build a startup dedicated entirely to world models.
Before that, Google’s Genie 3 had already sent shockwaves through the industry.
Though all three giants are pushing into world models, each represents a completely different bet on the future—and a radically different technical path.
The battle of world models has begun.
Fei-Fei Li had just published a lengthy manifesto championing spatial intelligence when her startup World Labs quickly followed up by launching Marble, its first commercial world model.
According to the team, this approach dramatically reduces scene distortion and detail inconsistency, while supporting export to Gaussian splats, mesh models, and even direct video output.
Taking it one step further, Marble includes a native AI world editor called Chisel. With a single prompt, users can reshape the world however they like.
For VR creators or game developers, the full pipeline—from one sentence, to a generated 3D world, to one-click export into Unity—feels almost magical.
Yet a machine learning engineer on Hacker News pointed out that, despite the branding, Marble looks more like a 3D rendering system than a true world model.
Isn’t this just a Gaussian Splat model? I’ve worked in AI for years, and I still can’t tell what exactly “world model” is supposed to mean.
A Reddit user put it even more bluntly:
Turning images into 3D scenes using Gaussian splatting, depth, and inpainting is cool, but this is just a 3D Gaussian pipeline—not a robot’s brain.
Gaussian splatting refers to a breakthrough 3D modeling technique that has swept the field in recent years.
It represents a scene as thousands of floating, colorful, fuzzy Gaussian blobs. These blobs are then “splatted” onto the screen, blending naturally into a rendered image.
Imagine each Gaussian as a tiny, semi-transparent bubble with soft edges drifting in 3D space.
On its own, one bubble is shapeless. But thousands of these bubbles, arranged and rendered from varying viewpoints, form a rich and realistic 3D scene.
This method avoids the heavy workflows of traditional photogrammetry. Though it sacrifices some precision, it offers incredible speed and ease of editing.
Marble is built squarely on this idea.
But that also means Marble is not what many imagined—a world model built for robotic training.
Marble does construct a full world, but what we see is simply the renderable view—what a camera or screen can show.
In other words, Marble captures what the world looks like, but not the underlying physics of how that world operates.
That’s fine for humans, but a robot needs causal structure, not just pixels.
For example, humans intuitively know a ball rolls downhill on a slope. But for robots, variables like mass, friction, and velocity are crucial—and Marble contains none of them.
Perhaps for this reason, Marble’s own blog talks a lot about “world models,” “Gaussian splats,” “meshes,” and “video exports,” yet barely mentions robotics at all.
Still, on the commercial front, Marble has a clear advantage.
Unlike abstract research-grade world models designed for embodied intelligence, Marble is a practical tool ready for game and VR pipelines today.
But this raises a troubling question: is the much-hyped “world model path to AGI” just a marketing illusion?
Absolutely not.
There are world models built for real robotic intelligence—most notably LeCun’s JEPA.
LeCun’s view of world models has little to do with 3D graphics. Instead, it draws from control theory and cognitive science.
These world models don’t generate beautiful images because you’re not meant to “look” at them.
Their purpose is to help robots think ahead—predict future states before acting.
JEPA embraces this philosophy.
LeCun argues that only the abstract intermediate representation matters. The model doesn’t need to waste compute generating pixels; it should focus entirely on world states that inform decision-making.
So while JEPA cannot produce the eye-catching 3D visuals that Marble can, it is far closer to a robot’s brain.
It learns the underlying structure of the world—making it an ideal training ground for embodied AI.
This creates a stark divide between the two visions of “world models”:
Fei-Fei Li’s approach resembles a front-end asset generator.
LeCun’s approach is more like a back-end predictive engine.
And right between these two legends stands another giant: Google.
In August, Google DeepMind launched Genie 3—its most advanced world model to date.
With a single prompt, Genie 3 can generate a fully interactive video environment that users can explore for several minutes.
Perhaps its most impressive leap is solving long-term scene consistency. No more turning around and watching buildings disappear.
It also supports world events like “start raining” or “nightfall,” creating the feeling of a game world governed not by a traditional engine but by the model itself.
However, Genie is best described as a “world-model-inspired video generator.”
Although Genie 3 animates the world, its core is still video generation, not the physics-based causality that powers JEPA.
It can be used for AI agent training, but it lacks JEPA’s deep physical grounding.
At the same time, its resolution and realism cannot match Marble’s high-precision, exportable 3D assets.
To summarize, all three world models describe “worlds”—but from completely different perspectives:
Marble shows what the world looks like.
Genie 3 shows how the world changes.
JEPA tries to understand why the world works the way it does.
Almost every world model on the market today fits into one of these three paradigms:
1. World Model as Interface
Represented by Marble, this approach generates visually rich, editable 3D environments for humans.
2. World Model as Simulator
Represented by Genie 3, this category produces dynamic, steerable environments where agents can learn by trial and error.
3. World Model as Cognitive Framework
Represented by JEPA, this focuses on latent representations and state transitions—the ideal training environment for robots.
According to researcher Zhao Hao, these can be stacked into a “world-model pyramid,” from Fei-Fei Li at the bottom, to Genie 3 in the middle, to LeCun at the top.
Looking up this pyramid:
The higher you go, the more abstract and aligned the model becomes with AI cognition and robotic training.
The lower you go, the more realistic and human-friendly the visuals—but the less meaningful they are to robots.
The world-model race has only just begun.
Which path leads closest to AGI remains an open question—but the competition between Fei-Fei Li, LeCun, and Google is already shaping the future of intelligent systems.
Go Master Showdown: Defending Champion Ke Jie Defeats Tu Xiaoyu to Seize Match Point in the Finals
Ke Jie Seizes Opening Victory in Go Sage Championship Final, One Step Away from Title Defense Septem...
Tiffany Woo × Cai Peixuan Make Their Web Drama Debut in “Battle 1212” — Dual Female Leads Ignite the Fiercest Workplace Showdown
Genting Highlands, 20 — The web series Battle 1212, written and directed by Monsterz, has officially...
Man United Eyes Four Consecutive Wins Ahead of Derby! Wolves’ Key Player Suspended, Amrolín May Revert to 3-4-3
On Tuesday evening, Manchester United will host bottom-placed Wolves in the Premier League, kicking ...
“Best gift for my 30th anniversary in the industry” — Daniel Chan finally collects his Bauhinia Best Actor trophy
(Hong Kong, Jan 9) Hong Kong actor-singer Daniel Chan Hiu-tung recently won Best Actor in the Main C...
Report: Honor Magic8 RSR May Be the Last Micro-Curved Display Flagship Among Top-Tier Smartphone Makers
IT Home reported on January 25 that blogger @體驗 more shared an update suggesting the Honor Magic8 RS...
Is Zhang Yuqi Losing Another Job? A Tour Guide Posts Photos in Support, While Yu Shi’s Fans Issue “Harsh Warnings,” Sparking Buzz
In recent weeks, Zhang Yuqi has once again found herself at the center of online controversy. A wave...
China’s 5th Gold! Eileen Gu Defends U-Shaped Freestyle Skiing Title, Becomes History-Making Athlete
In the women’s U-shaped freestyle skiing final at the Winter Olympics, Eileen Gu delivered a flawles...
Back-and-Forth Battle: Cavaliers Hold 12-Point Lead Before Knicks Rally, Mitchell vs. Brunson, Harden Adds 9 Points and 4 Assists
The NBA regular season continued with the Cavaliers hosting the Knicks. Both teams s...
Zhu Yawen spotted at a theme park with wife Shen Jiani and their two daughters — dressed more simply than passersby, daughters’ slim figures draw attention
During the Qingming holiday, a passerby spottedZhu Yawen and his wife Shen Jiani at Shanghai LEGOLAN...
Sun Yang and Zhang Doudou enjoy a sweet day at Disneyland, posing playfully with Daisy Duck with bright smiles
Sun Yang and his wife Zhang Doudou recently shared sweet moments from their Disney trip on social me...
Wins 460,000 in prize money! Zhao Xintong beats Ding Junhui 13-9 to reach quarterfinals, opponent confirmed
On April 26, Zhao Xintong defeated Ding Junhui 13–9 in a highly anticipated all-Chinese clash at the...