← Back to feed
2026-08-04agentsreasoningalignmentcode

A game theory for foundation models shows new paths to rational cooperation through similarity inference

Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, João Sacramento, Blaise Agüera y Arcas

PDF preview for A game theory for foundation models shows new paths to rational cooperation through similarity inference
Read on arXiv →

Key claim

Embedded agency leads to stable cooperation among AI agents.

In plain English

Imagine you're developing AI agents that need to work together in complex social situations, like negotiating deals or collaborating on tasks. Traditional game theory assumes that each agent acts independently, which often leads to suboptimal outcomes, like mutual defection in dilemmas. This approach breaks down when agents are designed to predict their own actions while considering the behavior of others, leading to unexpected results in cooperative scenarios. This is what's called decoupled agency, where agents don't account for their influence on one another. In contrast, the new approach introduces the concept of embedded agency, where agents view themselves as part of the environment they operate in. By doing so, they can infer the likelihood of cooperation based on their own decisions and the behavior of similar agents. This shift allows for a new equilibrium concept, termed embedded equilibrium, which better captures the dynamics of modern AI interactions. For builders, this means that designing agents with an understanding of their interconnectedness can lead to more effective cooperation strategies, moving beyond the limitations of classical models.

Novelty
8.5/10

Introduces a new theoretical model for AI agents that challenges classical game theory.

Reliability
7.5/10

Provides empirical findings in stylized social dilemmas, though further validation may be needed.

Deep reliability assessment

The methodology supports the claim that embedded Bayesian agents can achieve stable cooperation through similarity inference, but the generalizability to real-world scenarios may be overclaimed without empirical validation beyond stylized environments.

Reproducibility

yes, the complete experimental code is available at github.com/paradigms-of-intelligence/game-theory-foundation-models

Key figure

Figure 1 illustrates the architecture of rational foundation model agents integrating predictive models with optimal planning.

GitHub1 repo
paradigms-of-intelligence/game-theory-foundation-modelsOfficial