AI's Models of the World, and Ours

Jon Kleinberg

International Conference on Machine Learning 2025 · Invited Talk

Overview

Jon Kleinberg, a renowned professor from Cornell University, delivered a thought-provoking talk at ICML 2025, delving into the profound implications of AI's burgeoning capabilities, particularly its development of "models of the world." The presentation explored how machine learning, having long surpassed human performance in domains like chess, is now grappling with challenges ranging from the emergence of algorithmic monoculture to the intricate dynamics of human-AI collaboration and the fundamental theoretical underpinnings of language generation. Kleinberg emphasized that understanding these new frontiers requires not only empirical investigation but also a return to foundational computational theory, offering abstract models to reason about AI's behavior and its impact on human endeavors.

Watch on SlidesLive

Visual summary for AI's Models of the World, and Ours by Jon Kleinberg
Visual summary for AI's Models of the World, and Ours by Jon Kleinberg

Key moments

  1. 0:00 Introduction: ICML's growth and current AI challenges
  2. 2:00 Seeking domains with existing superhuman AI capabilities
  3. 3:00 Chess as a 'post-singularity' creative human pursuit
  4. 4:00 Chess: The 'Drosophila of superhuman AI'
  5. 5:00 Deep Blue vs. Kasparov and AI's rating dominance
  6. 6:00 AI inverts knowledge: Spectators now know more than players
  7. 7:00 Shift from chess aesthetics to pure algorithmic calculation

AI's Models of the World, and Ours

Speakers: Jon Kleinberg

Conference: ICML 2025

YouTube: https://slideslive.com/39043345

Overview

Jon Kleinberg, a renowned professor from Cornell University, delivered a thought-provoking talk at ICML 2025, delving into the profound implications of AI's burgeoning capabilities, particularly its development of "models of the world." The presentation explored how machine learning, having long surpassed human performance in domains like chess, is now grappling with challenges ranging from the emergence of algorithmic monoculture to the intricate dynamics of human-AI collaboration and the fundamental theoretical underpinnings of language generation. Kleinberg emphasized that understanding these new frontiers requires not only empirical investigation but also a return to foundational computational theory, offering abstract models to reason about AI's behavior and its impact on human endeavors.

The talk articulated a compelling narrative, beginning with the historical context of AI's engagement with chess as a "Drosophila" for intelligence, and then extending these insights to the broader landscape of modern AI systems, especially large language models (LLMs). Kleinberg highlighted the breakdown of human intuition in the face of superhuman AI, necessitating new frameworks for analysis. He introduced novel concepts such as mimetic models designed to emulate human behavior, and the critical validity-breadth tradeoff inherent in language generation.

This exploration is crucial for the AI/ML community as it navigates an era where AI is not merely a tool but an increasingly autonomous agent shaping our world. Kleinberg's work provides a lens through which to examine the societal, ethical, and practical challenges posed by powerful algorithms, urging researchers to develop more robust and interpretable AI systems. The talk’s synthesis of theoretical computer science, behavioral economics, and practical machine learning offers a comprehensive perspective on building AI that is not only capable but also aligns with human values and comprehension.

Background

▶ Watch: Introduction: ICML's growth and current AI challenges (0:00)

The field of machine learning, and by extension AI, has experienced an "unprecedented growth" throughout its history, a theme recurrent even in the early proceedings of conferences like ICML, dating back to 1983. This growth has culminated in the current era where AI capabilities often exceed human intuition, prompting a re-evaluation of how we understand and interact with these systems.

A central theme in AI's history, and a cornerstone of Kleinberg's talk, is chess. From the early musings of Alan Turing and Claude Shannon to the pioneering work of Newell and Simon, chess has served as a benchmark for artificial intelligence. John McCarthy famously dubbed it the "Drosophila of AI," a simple yet complex domain ideal for studying intelligence. The advent of Deep Blue in 1997, which famously defeated Garry Kasparov, marked a pivotal moment, signaling AI's ability to compete at the highest human levels. However, as Kleinberg pointed out, this was merely a coincidence of timing; AI chess engines continued to improve relentlessly. By 2005, human players could no longer consistently defeat desktop-level chess engines. Today, top human players (rated around 2800 ELO) are significantly outmatched by engines (rated 3500-3600 ELO), placing chess firmly in a "post-singularity" world where AI reigns supreme.

This superhuman performance in chess has led to several interesting observations. Spectators, armed with powerful engines, often understand a game's optimal lines better than the grandmasters themselves. The aesthetics of chess play have shifted from elegant, human-preferred positions to tactically brutal, engine-derived moves, regardless of visual appeal. Crucially, AI engines, when used for training, often "set humans up to fail" – they suggest brilliant, deeply tactical lines that, if deviated from slightly, leave the human player lost without the AI's continued guidance. This highlights a fundamental challenge: powerful AI's actions can be incomprehensible to humans, even if optimal.

Beyond chess, the talk contextualized the rise of powerful algorithms within broader societal concerns. The proliferation of a few dominant AI models, what Kleinberg terms algorithmic monoculture, poses risks in high-stakes domains where diverse perspectives and "second opinions" are critical. This problem underscores the need for methods to create diverse AI behaviors or, alternatively, to build models that explicitly mimic human-like decision-making, leading to the concept of mimetic models.

Finally, the talk grounded these observations in theoretical computer science, particularly in the realm of language learning. Decades ago, Mark Gold's theorem (1967) established the impossibility of language identification (determining the true language from positive examples alone) in the limit without strong assumptions. This historical result has shaped much of our understanding of learning. Kleinberg's work, however, revisits this space to pose a distinct, yet equally fundamental question: is language generation possible under similarly minimal assumptions? This theoretical inquiry provides a crucial backdrop for understanding the capabilities and limitations of modern LLMs, which are, at their core, sophisticated sequence generators.

Key Findings

▶ Watch: Chess as a 'post-singularity' creative human pursuit (3:00)

Kleinberg's talk presented several critical findings and conceptual contributions, bridging empirical observations from chess with deep theoretical insights into AI capabilities:

  1. Chess as the "Drosophila of Superhuman AI": The talk established chess as a prime example of a creative human domain where AI has not only surpassed human capabilities but has done so to an extreme degree. Modern chess engines operate at an ELO rating of 3500-3600, significantly higher than the best human players (around 2800). This "post-singularity" state in chess offers a valuable microcosm for studying the broader implications of superhuman AI across various domains.
  1. Algorithmic Monoculture vs. Mimetic Models: Kleinberg highlighted a spectrum of AI model deployment. On one end is algorithmic monoculture, where a few highly powerful algorithms dominate, leading to a lack of diversity in solutions and making "second opinions" difficult, particularly in high-stakes applications. On the other end are mimetic models, such as Maia, which are specifically designed to replicate human behavior, rather than simply optimizing for performance. This distinction raises important questions about diversity, interpretability, and the ethical implications of AI.
  1. AI's Tendency to "Set You Up to Fail": A key observation from human-AI interaction, particularly in chess, is that powerful AIs can make brilliant, optimal moves that are, however, incomprehensible or difficult for humans (or weaker AIs) to follow up on. When the AI partner is no longer present, the human is left in a difficult position. This phenomenon underscores a challenge in human-AI collaboration and points to the need for AIs that are "comprehensible" partners.
  1. The Efficacy of "Partner Bots": Counterintuitively, a weaker AI can be a more effective partner for another AI than a stronger one. Experiments demonstrated that a specially trained partner bot, despite being significantly weaker than a top engine like Leela, performed better when collaborating with Maia against a strong opponent. This is because its moves were more "comprehensible" and allowed Maia to follow up more successfully, leading to a higher win rate (approximately two-thirds of games). This finding suggests a paradigm shift in designing collaborative AI.
  1. Probing for Internal World Models in Sequence Generators: For pure sequence generation models like LLMs, which simply output text or moves, there's no explicit requirement for an internal "world state." Kleinberg emphasized the importance of using formal tools, like principles derived from the Myhill-Nerode theorem, to probe for and quantify the presence and accuracy of such implicit internal representations. This is crucial for understanding how LLMs reason and predict, especially in structured environments.
  1. Language Generation is Fundamentally Different from Language Identification: Building on Gold's theorem (1967), which shows the impossibility of language identification in the limit, Kleinberg presented a new theorem (with Sendhil Mullainathan) demonstrating that language generation is possible in the limit, even for arbitrary formal languages and with minimal assumptions. This theoretical breakthrough provides a new foundation for understanding the capabilities of LLMs as generators.
  1. The Inherent Validity-Breadth Tradeoff in Language Generation: The algorithm proposed for language generation in the limit, while guaranteeing validity (no hallucinations), inherently suffers from low breadth, meaning it generates a very sparse subset of the true language. This highlights a fundamental validity-breadth tradeoff, where perfect correctness can come at the cost of diversity or richness of output. Subsequent work, however, indicates that this damage can be limited, allowing for the generation of a positive fraction of the language. This mirrors the "linguistic aging" observed in human communities, where initial invalidity gives way to rigid, narrow expression.

Technical Deep Dive

▶ Watch: Chess: The 'Drosophila of superhuman AI' (4:00)

Kleinberg's talk traversed several technical dimensions, from specific AI system designs to abstract computational theory.

The journey began with Maia, a chess engine developed in collaboration with Ashton Anderson. Unlike traditional engines like Stockfish or Leela (an open-source implementation of AlphaZero), whose objective is to maximize win probability, Maia's goal is to match human moves. This is achieved by "ripping the head off" Leela's neural network and retraining it on a massive dataset of hundreds of millions of public human games from Lichess. Multiple versions of Maia are trained, each targeting a specific human ELO rating (e.g., Maia 1100, Maia 1200, up to Maia 2000). Maia always plays the arg max move – the move it predicts is most likely to be played by a human at its target rating. A key observation was that Maia 1100, when playing its most likely moves, performs at a higher rating (around 1400 ELO), suggesting that human players at lower ratings are often hampered by idiosyncratic errors. Recent work, such as the Ally chess engine from CMU, has extended this move-matching performance to grandmaster levels (2500-2600 ELO), addressing Maia's limitation of sparse training data at higher human skill tiers.

The concept of algorithmic monoculture was introduced to describe scenarios where powerful algorithms, due to their complexity and development cost, lead to a limited diversity of solutions. This contrasts with mimetic models, which are designed to mimic specific individual behaviors. For instance, one could train a Maia-like model on a specific player's 1000-2000 games to create a "signature" model of their play. This opens avenues for historical simulations (e.g., a Fischer bot vs. a Karpov bot) but also raises complex normative questions about identity, reputation, and harm when a digital "mimetic model" of a person makes mistakes or exhibits certain behaviors.

The issue of AI "setting you up to fail" was explored through the lens of human-AI or AI-AI collaboration. When a powerful AI like Leela makes a brilliant, non-obvious move, a human or a less sophisticated AI partner might not understand the follow-up, leading to a loss. To address this, Kleinberg's team developed partner bots. These bots are trained not just to play well, but to play as a partner for another AI (e.g., Maia), making moves that Maia can effectively follow up on. In experiments, a partner bot, while individually weaker than Leela (scoring 0.14 against it), proved to be a more effective partner for Maia, winning approximately two-thirds of games against a Leela-Maia opposition. This demonstrates the importance of comprehensibility and collaborative intelligence over raw individual strength.

The discussion then shifted to world models within sequence generation models, particularly LLMs. The core question is whether a model generating sequences (e.g., chess moves, Othello moves, or narrative text) implicitly maintains an internal representation of the "state" of the world. Since LLMs are pure sequence generators, there's no explicit requirement for such a state. To probe this, Kleinberg utilized principles from the Myhill-Nerode theorem. This theorem, originally from formal language theory, states that two sequences are equivalent if and only if they cannot be distinguished by any suffix. By applying this to game sequences (e.g., a number-choosing game analogous to tic-tac-toe on a magic square), one can quantify the rate of Myhill-Nerode deviations as a measure of how far an internal representation deviates from a true state machine. This method helps assess if an LLM has a consistent internal "board state" or "narrative state." The analogy to computer graphics (e.g., rendering Gollum with physically accurate hair and skin properties to overcome the uncanny valley) further emphasized that explicit, accurate world models can be crucial for achieving believable and robust outputs.

Finally, Kleinberg introduced groundbreaking theoretical work on language generation as an abstract primitive. He distinguished it from language identification, which Gold's theorem (1967) proved impossible in the limit without strong assumptions. Kleinberg and Sendhil Mullainathan showed that language generation is possible in the limit. Their algorithm operates by maintaining a set of "consistent" languages (those that contain all observed training strings) and iteratively identifying "critical" languages – consistent languages that are also "thin" subsets of all earlier consistent languages. The algorithm's strategy is to generate from the "rightmost critical language" (the thinnest consistent language it has encountered). This ensures that eventually, it will generate strings from the true, secret language, even without ever "identifying" it.

This theoretical framework reveals an inherent validity-breadth tradeoff. To guarantee perfect validity (no "hallucinations"), the algorithm must become increasingly conservative, generating from a sparser and sparser subset of the true language (low breadth). This leads to a phenomenon akin to "linguistic aging," where generative models, like humans in online communities, initially make errors (invalidity) but then become rigid and narrow in their output to ensure correctness (mode collapse). Subsequent work by Fan Wei and others aims to mitigate this by developing more powerful algorithms that can generate a positive fraction of strings from the language while maintaining validity, thereby limiting the damage of this tradeoff.

Experimental Setup & Results

▶ Watch: AI inverts knowledge: Spectators now know more than players (6:00)

The talk, while heavily theoretical in its latter half, presented several key empirical observations and results, particularly from the domain of chess:

Maia Chess Engine:

  • Training Data: Maia was trained on hundreds of millions of public human chess games from Lichess. This vast dataset allowed for the statistical modeling of human move choices.
  • Objective: Unlike traditional engines, Maia's training objective was move-matching performance against human players, not winning games.
  • Versions: Multiple versions of Maia were created, each tuned to match the playstyle of a specific human ELO rating range (e.g., Maia 1100, Maia 1200, up to Maia 2000).
  • Results: Each Maia version demonstrated peak move-matching performance at its intended rating level, with performance gradually decreasing for ratings further away. A notable finding was that Maia 1100, when consistently playing its most likely move, exhibited an effective playing strength of approximately 1400 ELO. This suggests that low-rated human players' actual skill is higher than their ELO, but it's masked by occasional, idiosyncratic errors.
  • Limitations & Improvements: Maia's performance for grandmaster-level play (2500+ ELO) was limited due to the sparsity of high-level human game data. This has been addressed by more recent projects like CMU's Ally chess engine, which has shown improved move-matching capabilities extending into the grandmaster range.

Partner Bot Experiment:

  • Setup: This experiment involved a collaborative scenario where a "partner bot" was trained to play alongside Maia against a strong opponent, specifically a Leela-Maia team (where Leela and Maia alternated moves). The partner bot's objective was to enable Maia to win.
  • Baselines: The partner bot was evaluated against Leela playing alone.
  • Results: While the partner bot was significantly weaker than Leela when playing independently (achieving a score of only 0.14 against Leela, where 0.5 is equal strength), it proved to be a much more effective partner for Maia. When collaborating with Maia, the partner bot-Maia team won approximately two-thirds of their games against the Leela-Maia opposition. This demonstrates that an AI designed for "comprehensible" collaboration, even if individually weaker, can lead to superior team performance.

Myhill-Nerode Principles for Internal State:

  • Setup: For sequence generation models, the talk discussed using Myhill-Nerode principles to probe for the presence and consistency of an internal "board state" or "world model." This involves analyzing sequences of moves (e.g., in a number-choosing game or Othello) to see if two sequences that should lead to the same internal state are indeed indistinguishable by subsequent play.
  • Metrics: The primary "result" here is the concept of quantifying Myhill-Nerode deviations, providing a measure of how far a model's implicit state deviates from a true, ground-truth state machine. This is a theoretical framework for analysis rather than a specific reported number from an experiment in the talk itself, though the methodology is applicable to models playing structured games or generating narratives.

Language Generation Theorem:

  • Setup: The core of this finding is a theoretical proof, not an experimental setup. It describes an algorithm that, given a countable collection of languages and a sequence of positive examples from a secret language, can generate valid strings from that secret language in the limit.
  • Results: The theorem proves that language generation is possible under minimal assumptions, a stark contrast to the impossibility of language identification (Gold's theorem). However, the initial algorithm guarantees validity at the expense of breadth, generating a very sparse subset of the language.
  • Ongoing Work: Subsequent research (e.g., by Fan Wei and others) aims to improve this, demonstrating that it's possible to generate a positive fraction of the language while maintaining validity. The current theoretical lower bound for this positive fraction is 1/8, with an upper bound of 1/2, leaving an open question for future research.

Practical Implications

▶ Watch: Shift from chess aesthetics to pure algorithmic calculation (7:00)

The insights from Jon Kleinberg's talk carry significant practical implications for practitioners, infrastructure teams, model builders, and deployers in the AI/ML space:

  1. Navigating Superhuman AI in High-Stakes Domains: The "post-singularity" reality in chess serves as a powerful analogy for other critical areas. As AI surpasses human capabilities in fields like medical diagnosis, legal reasoning, or financial trading, practitioners must move beyond simply comparing AI to human performance. The focus shifts to understanding AI's unique modes of operation, its potential for incomprehensible "brilliant" moves, and how to integrate its outputs safely and effectively into human workflows.
  1. Mitigating Algorithmic Monoculture: The risk of relying on a few dominant, complex AI models is significant, especially in high-stakes decision-making. Infrastructure teams and model builders should prioritize diversity in model architectures, training data, and objective functions. This could involve fostering open-source alternatives, encouraging varied research approaches, or developing ensemble methods that leverage diverse AI perspectives to avoid systemic biases or catastrophic failures stemming from a single, flawed algorithmic approach. The call for "second opinions" from AI becomes paramount.
  1. Designing for Human-AI Collaboration: The observation that AI can "set humans up to fail" highlights a critical challenge in human-AI teaming. Model builders need to consider not just optimality but also comprehensibility and predictability when designing AI partners. This might involve training AIs with objectives that explicitly account for human follow-through, generating explanations for their decisions, or even intentionally playing "sub-optimal" but more human-understandable moves. The success of the "partner bot" suggests a paradigm shift: an AI's value in a collaborative setting might not solely depend on its individual strength but on its ability to enable its human or AI teammates.
  1. Ethical and Normative Considerations for Mimetic Models: The ability to create highly personalized "mimetic models" that emulate specific individuals (e.g., a "Bobby Fischer bot" or a bot mimicking a specific Lichess player) opens up powerful applications but also profound ethical questions. Deployers and model builders must grapple with issues of identity, consent, reputation, and potential harm. If a mimetic model of a person makes a mistake or behaves inappropriately, does that reflect on the real person? How do we ensure these models are used responsibly and without causing undue distress or misrepresentation?
  1. Enhancing LLM Interpretability and Reliability: The work on probing for internal world models using Myhill-Nerode principles is crucial for understanding how LLMs operate beyond mere sequence generation. For practitioners building and deploying LLMs, this implies a need for diagnostic tools that can assess whether an LLM maintains a consistent and accurate internal representation of the domain it's operating in. This could lead to more reliable LLMs, especially in structured tasks like coding, scientific reasoning, or complex data analysis, where an implicit "world model" is essential for correctness and avoiding "hallucinations."
  1. Managing the Validity-Breadth Tradeoff in Generative AI: The theoretical finding of an inherent tradeoff between ensuring perfect validity (no hallucinations) and generating a broad, diverse range of outputs has direct implications for LLM design. Model builders must consciously decide where to position their models on this spectrum. For critical applications (e.g., medical advice, legal documents), absolute validity might be prioritized, even if it means sacrificing some creativity or diversity (leading to "mode collapse" or "linguistic aging"). For creative applications (e.g., story writing, artistic generation), a higher tolerance for occasional "hallucinations" might be acceptable to achieve greater breadth and originality. Future research aiming to achieve positive density generation while maintaining validity is a key area for practical advancement.
  1. Rethinking Abstraction Barriers in AI Systems: Kleinberg's "Alice and Bob" metaphor for language generation as an abstraction barrier challenges the community to define simpler, more fundamental contracts for AI's core capabilities. This could lead to more modular and robust AI system designs, where complex underlying models interact through well-defined interfaces, making them easier to reason about, debug, and improve.

Key Takeaways

  • Chess as a Bellwether for Superhuman AI: Chess provides a clear, long-standing example of a domain where AI has vastly surpassed human capabilities, offering crucial lessons for understanding the implications of powerful AI in other high-stakes fields.
  • The Dual Challenge of Algorithmic Monoculture and Mimetic Models: The prevalence of a few dominant AI solutions (monoculture) poses risks in critical applications, while the rise of personalized "mimetic models" designed to emulate human behavior introduces complex ethical and normative questions.
  • Designing AI for Comprehensible Collaboration: Powerful AIs can "set humans up to fail" by making brilliant but incomprehensible moves. Future AI systems must prioritize making "comprehensible" moves to facilitate effective human-AI or AI-AI collaboration, as demonstrated by the "partner bot" concept.
  • Probing for Internal World Models in LLMs: For pure sequence generation models like LLMs, it's vital to develop methods (e.g., using Myhill-Nerode principles) to infer and quantify the presence and accuracy of implicit internal "world models" to enhance interpretability and reliability.
  • Language Generation is Fundamentally Possible (Theoretically): A new theorem shows that language generation, unlike language identification, is possible in the limit under minimal assumptions, providing a theoretical foundation for understanding LLM capabilities.
  • The Inherent Validity-Breadth Tradeoff: There is a fundamental tension in generative AI between ensuring perfect validity (no hallucinations) and generating a broad, diverse range of outputs. This "linguistic aging" phenomenon requires careful consideration in AI design based on application needs.

About the Speaker(s)

Jon Kleinberg is a highly distinguished Professor of Computer Science at Cornell University, known for his interdisciplinary research spanning computer science, economics, and social systems. His work often explores the interface of algorithms and the real world, focusing on areas like network science, social networks, and the theoretical foundations of machine learning and AI. Kleinberg is a recipient of numerous prestigious awards, including the MacArthur Fellowship and the Nevanlinna Prize, recognizing his profound contributions to computational theory and its applications. His expertise allows him to bridge the gap between abstract theoretical concepts and their practical implications for rapidly evolving AI technologies.

Reviews

Maya Iyer (Theoretical ML Researcher) — SOLID

Kleinberg delivers a coherent and intellectually ambitious talk that weaves together empirical observations from chess AI, collaborative system design, and a genuinely interesting theoretical result about language generation. The crown jewel is the theorem with Mullainathan showing generation is achievable in the limit where identification (Gold 1967) is not — that is a clean, non-trivial separation between two fundamental primitives, and it's the kind of result the theory community should pay attention to. The chess material and partner-bot experiments are engaging but function more as extended motivation than rigorous empirical science; the Myhill-Nerode framing for probing world models…

Chen Zhao (Applied ML Researcher & Empiricist) — SOLID

Kleinberg's ICML 2025 talk is a well-constructed survey of ideas from his group's work on human-AI interaction, algorithmic monoculture, and the theoretical foundations of language generation. The theoretical contribution — distinguishing language generation from language identification and proving generation is possible in the limit — is genuinely interesting and underappreciated by the empirical ML community. The chess-derived empirics on Maia and the partner-bot result are clean and intuitive. But as a conference talk this is primarily a synthesis and framing piece, not a primary empirical contribution: the key experiments are previously published, the theoretical results are presented…

→ Top-rated talks at International Conference on Machine Learning 2025

All talks from International Conference on Machine Learning 2025