Publication:

Secret Robot: Evaluating Persuasion Capabilities of LLMs in the Social Deception Game of Avalon

Loading...
Thumbnail Image

Files

written_final_report.pdf (19.01 MB)

Date

2026-04-28

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

The opportunity for human-AI interaction has increased as AI continues to become more integrated into human society, raising concerns about the safety of AI systems and tools. The purpose of this study is to explore the persuasive capabilities of AI when deception is incentivized. Prior work has explored the potential of Large Language Models (LLMs) to reason strategically, participate in multi-agent environments, and human-AI interaction. Recent studies have compared how different LLMs perform in social deception games, but the idea of LLMs playing against humans in social deception games has not been explored extensively. Using the social deception game of Avalon, we explore the success rate of Evil teams made up of different compositions of AI agents and humans, comparing their performance against fully-human Evil teams. We develop a web-based implementation of Avalon that supports LLM-powered AI agents and run a human subjects study in which participants play games of Avalon across three randomized conditions that vary the number of AI agents on the Evil team. We find that Evil teams with AI agents, particularly the mixed team with a human and an AI agent on the Evil team, grossly underperform fully human Evil teams. Furthermore, we evaluate in-game data such as chat logs, voting behavior, and timing to understand mechanisms behind the Evil team success rates and to determine what persuasive strategies AI agents employ in Avalon.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation