Publication: One Step to Alignment: Pragmatic–Pedagogic Assistance Games
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
This thesis identifies and overcomes fundamental alignment limitations inherent to traditional inverse learning approaches. In particular, classical methods model the human as acting in isolation and the robot as a passive observer. However, this assumption imposes a hard bound on what the robot can infer, no matter how rational the human is. We prove this ceiling and then break it. By recasting the interaction as an \textit{assistance game}, where the human behaves pedagogically and the robot reasons pragmatically, we show that a single action can fully disambiguate the human's goal within just one time step. We formalize the structural conditions under which this holds, introduce a tractable best-response algorithm to compute optimal assistance strategies, and validate our theoretical results in a collaborative building domain. The result is a robot assistant that is corrigible by design: responsive to correction and one step from alignment.