Philosophical Transactions of the Royal Society A · Volume 384 · Issue 2320 · 2026

World Models

in natural and artificial intelligence — eighteen papers, one running argument.

18 papers · 217 concepts · 28 themes · 15 tensions · 84 quoted positions
Bradly AliceaNadav AmirRuairidh M. BattledayFritz BreithauptKatherine M. CollinsEunice YiuDouglas HofstadterDavid C. KrakauerAlexander Y. KuEvelina LeivadaBenjamin LyonsGeorg NorthoffVickram N. PremakumarNicolas RouleauFrancesco SaccoAdam SafronPedro TsividisHanlin Zhu
the eighteen papers in embedding space

In this issue

  1. null
    World models, artificial general intelligence and the hard problems of life-mind continuity: toward a unified understanding of natural and artificial intelligence
    Adam Safron, Michael Levin, Victoria Klimaj, Zahra Sheikhbahaee, Dalton Sakthivadivel, Adeel Razi, David Ha, Nick Hay, Kevin Schmidt, Irina Rish, David Krakauer, Melanie Mitchell, Samuel J. Gershman, Joshua B. Tenenbaum
    editorial overview cognitive sciencephilosophy of mindartificial intelligencemachine learningtheoretical biologycomplex systemsneuroscience

    This editorial introduces the special issue by asking what a 'world model' is and to what extent current AI systems possess one. Refusing a single definition, it surveys 17 contributions spanning causal, self-referential, goal-directed, collective and narrative forms of world modelling in biological and artificial systems. Its synthesizing conclusion is that while recent AI advances are revolutionary, LLMs likely model the world through statistical surface regularities rather than sufficiently coherent causal, embodied and value-laden world models, and so remain far from human-like general intelligence, consciousness or artificial superintelligence.

  2. 01
    Is there an 'I' in AI?
    Douglas Hofstadter
    conceptual/position artificial intelligencecognitive sciencephilosophy of mind

    When words 'act like' things in the world—when a system's verbal behaviour matches or meshes with the world's behaviour—those words refer to and mean those things; behind meaningful words lie concepts, and where ideas are put together in ways that make sense, there is thinking, and thinking is consciousness with a genuine 'I'. By this criterion, today's LLMs have meanings and concepts to a limited degree, because accurate mirroring of the world comes in shades of grey, not as a zero-versus-one jump. No magical extra ingredient beyond world-matching verbal behaviour (essentially the Turing Test) is needed for membership in the 'Thinking Club', and computational systems with real 'I's—self-pointed strange loops—may arrive frighteningly soon, perhaps by 2033.

  3. 02
    Large language models and emergence: a complex systems perspective
    David C. Krakauer, John W. Krakauer, Melanie Mitchell
    conceptual/position complex systemsmachine learningcognitive sciencephilosophy of science

    Claims of 'emergence' in LLMs—usually based on sudden jumps in benchmark accuracy with scale or on abilities the models were not explicitly trained for—do not meet the scientific standard of emergence in complexity science, which minimally requires a coarse-graining of observables coupled to a compressed system description that remains predictive (an effective theory). The authors lay out a framework of five conditions for emergence (scaling, criticality, compression, novel bases, generalization) and distinguish 'knowledge-out' emergence in simple physical systems from 'knowledge-in' emergence in adaptive systems like LLMs, where behavioral evidence alone cannot establish emergence without identifying the internal micro-to-macro reorganization. They further distinguish emergent capability from emergent intelligence—the efficient, analogy-mediated use of coarse-grainings to solve broad problem ranges ('less is more')—arguing that LLMs at best demonstrate emergent capability, while emergent intelligence remains unproven in them.

  4. 03
    Levels of analysis for large language models
    Alexander Y. Ku, Declan Campbell, Xuechunzi Bai, Jiayi Geng, Ryan Liu, Raja Marjieh, R. Thomas McCoy, Andrew Nam, Ilia Sucholutsky, Veniamin Veselovsky, Liyi Zhang, Jian-Qiao Zhu, Thomas L. Griffiths
    review/synthesis cognitive sciencemachine learningcognitive psychologyneuroscience

    Modern AI systems such as large language models are increasingly powerful but increasingly hard to understand, a predicament the authors argue is analogous to cognitive science's historical difficulty in understanding the human mind from the outside. They propose that methods refined over decades of cognitive science can be applied to LLMs, organized by Marr's three levels of analysis: at the computational level, the training objective and optimal benchmarks (Bayesian models, axiomatic systems) predict and diagnose behaviour; at the algorithmic level, behavioural methods from cognitive psychology (error patterns, similarity judgments, implicit associations) reveal representations and processes; and at the implementation level, neuroscience-inspired representational and causal analyses illuminate how these are realized in artificial neurons. This levels-based toolkit both explains LLM behaviour and enables adversarial identification of tasks that should be uniquely difficult for these non-human cognitive architectures.

  5. 04
    A sentence is worth a thousand pictures: can large language models understand human language?
    Evelina Leivada, Gary Marcus, Fritz Günther, Elliot Murphy
    empirical study linguisticscognitive sciencemachine learning

    Large language models are useful tools but not faithful models of human language: claims of human-like or suprahuman linguistic ability rest on weak evaluation standards and faulty inference from benchmark accuracy to understanding. Because LLMs lack grounded cognition, they cannot exploit top-down feedback from previous expectations and past world experience, relying instead on fixed associations between represented words and word vectors. A novel 'leet task' requiring decoding of sentences with letters replaced by numbers confirms this: humans perform near ceiling while models struggle, show accuracy-reasoning mismatches, and commit distinctly non-human errors. The missing abilities—grounding, generative world models, compositional structure—require solutions that go beyond increased system scaling.

  6. 05
    What does it mean to have an experience? Two kinds of narrative world models
    Fritz Breithaupt
    conceptual/position cognitive sciencenarratologyphilosophy of mindartificial intelligence

    Alongside traditional knowledge-based world models, agents build narrative world models that organize actions and events into temporal episodes with beginnings and endings linked to emotional outcomes. Within these, the paper proposes a distinct subset of experience-focused world models oriented toward novel, potentially transformative future experiences: before the event, radical uncertainty triggers intense mental activity and multiple contradictory imagined versions; afterwards, the agent may cast non-predetermined aesthetic, preference or moral judgments that can reshape the worldview and lead to self-updating. Such experience-focused world models are 'positively limited'—defined by limitation and exclusion rather than knowledge addition—and whether current AI could or should develop them remains an open question.

  7. 06
    Human-level learning of complex novel tasks as theory-based modelling
    Pedro Tsividis, João Loula, Jake Burga, Juan Pablo Rodriguez, Sergio Arnaud, Nate Foss, Andres Campero, Ajay Subramanian, Thomas Pouncy, Samuel J. Gershman, Joshua B. Tenenbaum
    computational modelling cognitive sciencemachine learningartificial intelligence

    Humans learn complex novel tasks such as new video games in just minutes, orders of magnitude faster than deep reinforcement learning systems, and this ability can be captured by a particularly strong form of model-based RL the authors call theory-based reinforcement learning, in which human-like intuitive theories — rich, abstract, causal models of physical objects, intentional agents and their interactions — guide exploration, model learning and planning. The authors instantiate this approach in EMPA (the Exploring, Modeling, and Planning Agent), which uses Bayesian inference to learn probabilistic generative models expressed as programs for a game-engine simulator and plans over internal simulations with theory-derived intrinsic rewards. EMPA matches human learning efficiency on a benchmark of 90 Atari-style games and reproduces fine-grained structure of human exploration, suggesting a way forward for more general, human-like AI.

  8. 07
    Unexpected benefits of self-modelling in neural systems
    Vickram N. Premakumar, Michael Vaiana, Florin Pop, Judd Rosenblatt, Diogo Schwerz de Lucena, Kirsten Ziman, Michael S. A. Graziano
    computational modelling machine learningneurosciencecognitive science

    When an artificial neural network is trained, as an auxiliary task, to predict its own internal states (self-modelling), it changes in a fundamental way: to better predict itself it restructures into a simpler, more regularized, more parameter-efficient and therefore more predictable system. The authors call this 'self-regularization through self-modelling' and demonstrate it empirically across several architectures and tasks. They argue it may explain observed benefits of self-models in machine learning and the adaptive value of self-models in biological brains, including making an agent more amenable to being modelled by others in social, cooperative settings.

  9. 08
    Empowerment gain and causal model construction: children and adults
    Eunice Yiu, Kelsey Allen, Shiry Ginosar, Alison Gopnik
    empirical study cognitive sciencedevelopmental psychologymachine learningphilosophy of science

    Empowerment—an intrinsic reward that maximizes the mutual information between an agent's actions and their outcomes—may be an important bridge between classical Bayesian causal learning and reinforcement learning. Because causation, on the interventionist view, just is the relation in which actions predictably produce outcomes, learning an accurate causal world model necessarily increases empowerment, and increasing empowerment leads to a more accurate (if implicit) causal world model. Empowerment may also explain distinctive features of children's causal learning and provide a more tractable computational account of how that learning is possible. Two experiments show that children and adults are sensitive to both controllability and variability—the two components of empowerment—when inferring causal relations, designing interventions and generalizing to new outcomes, objects and perceptual dimensions.

  10. 09
    Goals and the structure of experience
    Nadav Amir, Stas Tiomkin, Angela Langdon
    conceptual/position cognitive sciencereinforcement learningcomputational neurosciencephilosophy of mind

    Purposeful behavior is usually explained by world models split into a descriptive state representation and a prescriptive reward function that are treated as independent, with the descriptive enjoying epistemic primacy. The authors argue instead that the descriptive and prescriptive aspects of a world model co-emerge interdependently from an agent's goal. Drawing on Buddhist (Dharmakirti's) epistemology, they introduce 'telic states'—equivalence classes of goal-equivalent experience distributions—and formalize goal-directed learning as minimizing the statistical divergence between an agent's behavioral policy and desirable experience features. This yields a unified, goal-centered account of the behavioral, phenomenological, and neural dimensions of purposeful behavior across substrates.

  11. 10
    A 'good' regulator may provide a world model for intelligent systems
    Bradly Alicea, Morgan Hough, Amanda Nelson, Jesse Parent
    conceptual/position cyberneticsmachine learningcomplex systemscontrol theorycognitive science

    The paper recasts the classic cybernetic Every Good Regulator Theorem (EGRT) as a framework for building world models in intelligent autonomous learning systems, where a good regulator R forms a one-to-one, requisite-variety-matched mapping to a system S, and extends via second-order cybernetics to an internal model M that observes and supervises the S-R closed loop. By reframing physical phenomena such as temporal criticality, non-normal denoising and alternating procedural acquisition as statistical-mechanical regulatory relationships, the authors argue the EGRT can yield adaptive, physics-inspired world models that generalize to out-of-distribution and non-uniform task environments. At the same time these diverse physical systems expose the limits of tightly-coupled good regulation.

  12. 11
    Brains and where else? Mapping theories of consciousness to unconventional embodiments
    Nicolas Rouleau, Michael Levin
    conceptual/position philosophy of mindtheoretical biologydevelopmental biologycognitive scienceconsciousness science

    Most theories of consciousness (ToCs), once their neurocentric vocabulary is stripped away, describe functional principles that are not specific to neurons or brains. The authors argue that the mechanisms and algorithms found in brains are ancient and shared across cells, so minds may have preceded brains, and that the special status granted to brains reflects convention and the limits of human imagination rather than anything in the content of the theories. They conclude that the continuity of mind from single cells is the null hypothesis and that consciousness science should be open to minds in unconventional embodiments such as cells, tissues, organoids, biobots and life-technology hybrids.

  13. 12
    Cognitive glues are shared models of relative scarcities: the economics of collective intelligence
    Benjamin Lyons, Michael Levin
    conceptual/position theoretical biologyeconomicscomplex systemscognitive science

    The paper argues that the economy is a genuine collective intelligence whose 'cognitive glue' is the price system, which coordinates autonomous, self-interested agents by acting as a shared model of relative scarcities that lets each member form plans mutually compatible with everyone else's without central control. It then argues that any cognitive glue (for example, bioelectricity in bodies) must solve this same coordination problem in essentially the same way, so the price system is a generic abstract template for all cognitive glues. This points toward unifying biology and economics under a single theory of diverse/collective intelligence.

  14. 13
    What physics offers for artificial intelligence? Lessons from the brain's inner time and its dynamics
    Georg Northoff, Yasir Catal, Samira Abbasi
    conceptual/position neurosciencephilosophy of mindartificial intelligencecomputational neurosciencephysics

    Physics fundamentally offers time, and time is best understood as dynamics: the continuously changing patterns of activity unfolding over time. The brain's spontaneous activity generates an intrinsic temporal structure ('inner time')-manifested in scale-free dynamics and neural variability-that actively processes and encodes the world's input dynamics, allowing the organism to 'participate' in physical time and thereby to 'be in time' and 'be in the world.' Current computing devices, both classical and natural, lack spontaneous activity and inner time, so they encode inputs non-temporally and passively; they are 'locked out of time and world' and thus cannot acquire tacit knowledge or navigate a continuously changing world flexibly.

  15. 14
    On the representation complexity of model-based and model-free reinforcement learning
    Hanlin Zhu, Baihe Huang, Stuart Russell
    formal theory machine learningtheoretical computer science

    The paper studies why model-based reinforcement learning algorithms usually enjoy better sample complexity than model-free ones, proposing representation complexity, formalized via circuit complexity, as a key explanation. It proves that for a broad class of MDPs ('majority MDPs'), the underlying transition kernel and reward function can be computed by constant-depth, polynomial-size circuits, while the optimal Q-function requires exponential-size constant-depth circuits. Thus in many environments the ground-truth model of the world is simple to represent while derived quantities like Q-functions are complex, and experiments in MuJoCo environments corroborate that Q-functions are consistently harder for neural networks to approximate than transition and reward functions.

  16. 15
    Artificial intelligence for science: the easy and hard problems
    Ruairidh M. Battleday, Samuel J. Gershman
    conceptual/position cognitive sciencephilosophy of scienceartificial intelligencemachine learninghistory of science

    Recent AI-driven scientific breakthroughs solve only the 'easy problem' of science: optimizing a function whose inputs, outputs, objective, and data have already been specified in advance by human scientists. The 'hard problem' is coming up with the problem itself, which requires continual conceptual revision under poorly defined constraints and lies beyond current algorithms. Progress on the hard problem should begin with the cognitive science of how human scientists specify and respecify problems, and use those insights to build agents that autonomously infer and update their scientific paradigms.

  17. 16
    Topological constraints on self-organization in locally interacting systems
    Francesco Sacco, Dalton A. R. Sakthivadivel, Michael Levin
    formal theory theoretical biologycomplex systemsstatistical physicsmachine learningcognitive science

    All intelligence is collective intelligence, requiring parts to align toward system-level goals, and whether such a system can spontaneously reach and hold an ordered target state is a topological property of its interaction graph. By studying how free energy scales under the formation of domain walls in three model systems (the Potts model, autoregressive models, and hierarchical networks), the authors show that the combinatorics of local interactions determine when spontaneous ordering is possible. One-dimensional systems such as autoregressive language models cannot maintain long-range order, whereas hierarchical multiscale systems like biological tissues can, so topology is the critical factor distinguishing them.

  18. 17
    Revisiting Rogers' Paradox in the context of human-AI interaction
    Katherine M. Collins, Umang Bhatt, Ilia Sucholutsky
    computational modelling cognitive sciencemachine learninghuman-AI interactioncultural evolutioncomplex systems

    Rogers' Paradox showed that cheap social learning does not raise a population's equilibrium fitness above that of pure individual learning. The authors revisit this classic cultural-evolution result for an age in which humans increasingly socially learn from AI systems that are themselves socially learning from us. Extending Rogers' agent-based simulations with an AI node that learns from the whole population, they find the paradox re-emerges: cheap, widely available AI does not on its own improve a society's 'collective world model,' but strategies such as critical social learning, considered appraisal of when to trust the AI, and appropriate AI update schedules can shift the equilibrium.