Publication:

Training Large Language Models on Human Strategic Behavior

Loading...
Thumbnail Image

Files

written_final_report.pdf (1.02 MB)

Date

2026-04

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Training large language models to predict human behavior is a promising alternative to traditional cognitive models, but it remains unclear whether behavioral fine-tuning induces genuine strategic understanding or surface-level alignment. I tested twelve LLMs on a dataset of more than 2,400 two-player 2 × 2 matrix games, asking models to predict aggregate human choice distributions from natural language game descriptions. Larger models generally aligned better with human behavior, but none matched traditional cognitive models such as the level-k quantal response model. To improve alignment, I fine-tuned four instruction-tuned LLMs and evaluated them on both a primary task (predicting human choice distributions) and an out-of-distribution task (choosing actions in the games themselves rather than predicting human frequencies). Fine-tuning improved performance on both: models became more accurate at predicting human choice distributions, and their own choice probabilities became more aligned with Nash equilibrium play and empirical human behavior. However, in the out-of-distribution task, models reduced overconfidence uniformly rather than adapting to each game’s payoff structure. This demonstrates that behavioral alignment and context-dependent understanding are dissociable and that standard performance metrics may not distinguish between them.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation