Skip to content
KO EN
AI 기술 Upcoming

Can Language Models Spark Scientific Revolutions? DeepMind Researcher Says No—and Points to World Models

Google DeepMind researcher Tom Zahavy has dropped a provocative position paper titled "LLMs can't jump," arguing that current language models lack the cogn

Google DeepMind researcher Tom Zahavy has dropped a provocative position paper titled “LLMs can’t jump,” arguing that current language models lack the cognitive mechanism needed to drive scientific revolutions. Instead, he suggests that world models—systems that simulate and reason about the physical world—may be the path forward. The paper has stirred debate in AI circles, with implications that reach beyond academia into investment strategy and the future direction of the industry.

What Happened: Einstein’s Framework Reveals the Gap

Zahavy builds his case on a framework sketched by Albert Einstein in a letter to his friend Maurice Solovine. Discovery, Einstein wrote, is a cycle: sensory experience leads to an intuitive “leap” toward axioms, and from there logical deduction produces testable conclusions. To pinpoint where language models fall short, Zahavy draws on philosopher Charles Sanders Peirce’s triad of reasoning: deduction, induction, and abduction. Language models excel at deduction and induction—systems like AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could derive general relativity if given Einstein’s assumptions as a starting point. The bottleneck, he argues, lies in “manipulative abduction”: the creative leap that invents a cause for which no linguistic template yet exists.

Zahavy illustrates this with Einstein’s “happiest thought”: the freely falling observer who no longer feels gravity. This insight came from embodied simulation—Einstein mentally playing through a physical sensation, not grinding through equations. Similarly, Archimedes’ buoyancy principle emerged from the physical feeling of water rising as he stepped into a bathtub. Language models, Zahavy argues, lack this sensory grounding. He compares them to philosopher John Searle’s “Chinese Room” thought experiment: they shuffle symbols according to statistical rules without understanding.

Why It Matters: A Potential Pivot in AI Investment

This paper challenges the prevailing assumption that scaling up language models—more data, more compute—will eventually lead to human-level intelligence. If Zahavy is right, the current AI investment strategy, which has driven demand for NVIDIA (NVDA) GPUs and massive training runs, may need rethinking. Instead, capital could shift toward world models, reinforcement learning, and architectures that incorporate sensory feedback. Zahavy also highlights the “absence of error signal” problem: when Einstein was working, Newtonian physics was confirmed with extreme precision, and the only anomaly—a tiny shift in Mercury’s orbit—was attributed to a hypothetical planet. An optimization-driven AI would have had no reason to overthrow physics. This raises questions about whether data-centric AI is truly suited for scientific discovery in fields like drug development or materials science.

XPLAIN AI’s Interpretation: The Case for World Models

XPLAIN AI interprets this paper as signaling a necessary evolution beyond pure language models. Zahavy’s call for world models—agents that interact with environments and simulate outcomes—points toward AI systems capable of causal reasoning and physical intuition. This aligns with reinforcement learning and robotics. If research pivots in this direction, demand could grow for simulation platforms (e.g., NVIDIA Omniverse) and robotics hardware (e.g., Tesla TSLA, Boston Dynamics). Conversely, startups focused solely on scaling language models or optimizing inference costs may face technological headwinds. However, this remains a theoretical argument; experimental validation is still needed. Multimodal models like Gemini, which process text, images, and video, might partially bridge the sensory gap Zahavy identifies.

Potential Beneficiaries and Risks

  • Potential beneficiaries (🟢): Companies focused on world models, simulation software (e.g., Ansys), robotics (Tesla, Boston Dynamics), and AI research in reinforcement learning. Big tech firms like Google (GOOGL) and Microsoft (MSFT) could increase R&D in these areas.
  • Potential risks (🔴): Startups heavily dependent on pure language model scaling, or those optimizing inference costs for text-only models. Data labeling and collection firms might see reduced demand if simulation-based learning takes off.

These are inferences based on technical implications, not market predictions. The actual impact will depend on how the paper is received and what follow-up research shows.

Counterarguments and Uncertainty

Critics may argue that language models already learn implicit world knowledge from text, and that multimodal models are gaining sensory reasoning. Furthermore, Zahavy’s definition of manipulative abduction may be too narrow; many scientific discoveries arise from incremental knowledge and serendipitous experiments, not just intuitive leaps. The paper itself is a theoretical framework without concrete experimental results, so evidence that world models can perform manipulative abduction is still lacking. Investors should treat this as a signal, not a certainty.

What to Watch Next

Key indicators include: (1) follow-up experiments testing world models on novel hypothesis generation, (2) shifts in research funding from language models to world models, (3) announcements from major AI labs about new architectures, and (4) whether multimodal models begin to demonstrate creative abduction in controlled settings.

#DeepMind #WorldModels #AIResearch #ScientificDiscovery #LLMs #ReinforcementLearning #ArtificialIntelligence #TechInvesting

Sources

Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.

Found an error? Request a correction →