The Failure of Brute-Force Next-Token AGI

Our read
Feeding the entire internet into a giant transformer has reached its limit. True intelligence is the rate of skill acquisition on novel tasks, not memorizing Wikipedia. Without explicit world models that let an agent simulate actions and anticipate failure before moving, we are just building highly sophisticated, incredibly expensive parrots.
What happened
As frontier AI models consume the entire internet, researchers are realizing that static knowledge benchmarks do not equal true intelligence. The industry is beginning to split over whether to keep scaling autoregressive transformers or pivot to explicit world models that can simulate reality.
The brief
The ARC-AGI benchmark failures prove the point: when the pattern matches stop working, the current crop of LLMs fall flat on their face.
Key findings
The ARC-AGI benchmark failures prove the point: when the pattern matches stop working, the current crop of LLMs fall flat on their face.
Their pitch: Predicting the next token on larger datasets will eventually yield emergent reasoning and world-modeling capabilities.
The fight
Named sides below. The brief above already picked.
- Scale-is-All-You-Need Camp
Predicting the next token on larger datasets will eventually yield emergent reasoning and world-modeling capabilities.
- Model-Based Realists
Autoregressive LLMs are just high-fidelity memorizers that fail on novel tasks because they lack an explicit, sample-efficient physics engine.
Why now
Why now. The industry-wide discussion on the data wall, the plateauing of LLM performance, and François Chollet's ARC-AGI benchmark challenges.
From the episode. World Models, JEPA, and the Path to Sample-Efficient RL (https://www.youtube.com/watch?v=qz4GQ0zUFRw)
