Publication:

One Step to Alignment: Pragmatic–Pedagogic Assistance Games

Loading...
Thumbnail Image

Files

Lazarski_Elle_Thesis.pdf (1.41 MB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

This thesis identifies and overcomes fundamental alignment limitations inherent to traditional inverse learning approaches. In particular, classical methods model the human as acting in isolation and the robot as a passive observer. However, this assumption imposes a hard bound on what the robot can infer, no matter how rational the human is. We prove this ceiling and then break it. By recasting the interaction as an \textit{assistance game}, where the human behaves pedagogically and the robot reasons pragmatically, we show that a single action can fully disambiguate the human's goal within just one time step. We formalize the structural conditions under which this holds, introduce a tractable best-response algorithm to compute optimal assistance strategies, and validate our theoretical results in a collaborative building domain. The result is a robot assistant that is corrigible by design: responsive to correction and one step from alignment.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation