Publication: Training Large Language Models on Human Strategic Behavior
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Training large language models to predict human behavior is a promising alternative to traditional cognitive models, but it remains unclear whether behavioral fine-tuning induces genuine strategic understanding or surface-level alignment. I tested twelve LLMs on a dataset of more than 2,400 two-player 2 × 2 matrix games, asking models to predict aggregate human choice distributions from natural language game descriptions. Larger models generally aligned better with human behavior, but none matched traditional cognitive models such as the level-k quantal response model. To improve alignment, I fine-tuned four instruction-tuned LLMs and evaluated them on both a primary task (predicting human choice distributions) and an out-of-distribution task (choosing actions in the games themselves rather than predicting human frequencies). Fine-tuning improved performance on both: models became more accurate at predicting human choice distributions, and their own choice probabilities became more aligned with Nash equilibrium play and empirical human behavior. However, in the out-of-distribution task, models reduced overconfidence uniformly rather than adapting to each game’s payoff structure. This demonstrates that behavioral alignment and context-dependent understanding are dissociable and that standard performance metrics may not distinguish between them.