A game theory for foundation models shows new paths to rational cooperation through similarity inference
Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, João Sacramento, Blaise Agüera y Arcas
Read on arXiv →Key claim
Embedded agency leads to stable cooperation among AI agents.
In plain English
Imagine you're developing AI agents that need to work together in complex social situations, like negotiating deals or collaborating on tasks. Traditional game theory assumes that each agent acts independently, which often leads to suboptimal outcomes, like mutual defection in dilemmas. This approach breaks down when agents are designed to predict their own actions while considering the behavior of others, leading to unexpected results in cooperative scenarios. This is what's called decoupled agency, where agents don't account for their influence on one another. In contrast, the new approach introduces the concept of embedded agency, where agents view themselves as part of the environment they operate in. By doing so, they can infer the likelihood of cooperation based on their own decisions and the behavior of similar agents. This shift allows for a new equilibrium concept, termed embedded equilibrium, which better captures the dynamics of modern AI interactions. For builders, this means that designing agents with an understanding of their interconnectedness can lead to more effective cooperation strategies, moving beyond the limitations of classical models.
Introduces a new theoretical model for AI agents that challenges classical game theory.
Provides empirical findings in stylized social dilemmas, though further validation may be needed.
Deep reliability assessment
The methodology supports the claim that embedded Bayesian agents can achieve stable cooperation through similarity inference, but the generalizability to real-world scenarios may be overclaimed without empirical validation beyond stylized environments.
Reproducibility
yes, the complete experimental code is available at github.com/paradigms-of-intelligence/game-theory-foundation-models
Key figure
Figure 1 illustrates the architecture of rational foundation model agents integrating predictive models with optimal planning.
