The Limits of Next-Token Prediction
Recent generative AI successes have led some to believe that scaling multimodal models will inevitably lead to artificial general intelligence. However, this approach fundamentally misunderstands the nature of intelligence. Predicting the next token does not require a model of the physical world. Models like OthelloGPT demonstrate that systems can score remarkably well on sequence prediction by learning idiosyncratic heuristics rather than genuine world models.
Syntax Mimicry Versus Semantic Grounding
Human language understanding fuses syntax, semantics, and pragmatics. While large language models can generate syntactically well-formed sentences, they often fail at semantic and pragmatic reasoning because they lack grounded world knowledge. They reduce problems of meaning to brute-force memorization of abstract rules governing symbols, creating a superficial illusion of comprehension.
Embodiment as the Path Forward
Sutton's Bitter Lesson suggests that methods leveraging computational resources will outpace hand-coded structure. However, applying scale maximalism to AGI by summing narrow modalities is flawed. True general intelligence requires an interactive, embodied cognitive process where modality-specific processing emerges naturally from a unified perception and action system, rather than being artificially partitioned and glued together.