Same Game, Different Story: A Minimal Conservative Strategic Robustness Benchmark for Large Language Model Agents
This paper introduces a benchmark, Same Game, Different Story, to evaluate strategic robustness of large language model agents by measuring their action distribution invariance under payoff-preserving framing changes.
The paper introduces a new benchmark, Same Game, Different Story, to evaluate strategic robustness of large language models.
Before reading this…
Applications
- →Evaluating strategic robustness of large language models.
- →Understanding how language models behave in strategic settings.
To understand this paper, make sure you know these concepts first:
- Understanding of large language models.find papers →
- Familiarity with social dilemma games.find papers →
Abstract
More Like ThisLarge language model (LLM) agents increasingly operate in strategic settings where outcomes depend on the actions of other agents. This raises a reliability question: will a model choose consistently when the same incentives are presented through different narratives? We introduce Same Game, Different Story, a benchmark that defines strategic robustness as invariance of model-induced action distributions under payoff-preserving changes in framing. We illustrate the framework through a secondary analysis of published aggregate cooperation rates for GPT-3.5, GPT-4, and LLaMa-2 across four social-dilemma games. The retained comparison covers business and friend-sharing framings, representing 24 model-game-context cells and 7,200 decisions in the source study. Because trial-level data were unavailable, approximate counts were reconstructed from published figures; the resulting estimates are therefore illustrative rather than an exact replication. Under the paper's conservative transformation, pooled strategic robustness is 0.783, and friend-sharing framing increases cooperation by 0.307 relative to business framing. The results indicate that social-relational framing can substantially alter LLM behavior even when the underlying action sets and payoffs remain fixed. Strategic robustness should therefore be evaluated separately from strategic competence, using families of payoff-equivalent prompts rather than a single presentation of a game.