Publication: Agent Scaffolding and Self-Adaptive Methods for Embodied, Long-Horizon Speedrunning in Pokémon Emerald
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Advances in large language models have enabled the development of virtual agents that solve problems across practical, real-world applications. While these models have shown remarkable capability in optimizing for near-term progress, they struggle under the unique demands imposed by long-horizon settings. Few existing benchmarks emphasize partial observability, game-theoretic reasoning, and long-horizon planning simultaneously, leaving a gap between what evaluations measure and what real-world autonomy demands. Our work presents three contributions toward closing that gap. First, we introduce the PokéAgent Speedrunning Benchmark, a standardized evaluation framework set in Pokémon Emerald that requires agents to sustain coherent reasoning across thousands of interactions within a dynamic, partially observable environment. We further develop an expert agent harness that decomposes gameplay into a multi-agent orchestration system with purpose-built memory, planning, skills, subagents, and self- reflection capabilities. Tracing our expert harness through several stages of development, we identify key architectural decisions that most impact long-horizon performance. Finally, we introduce AutoEvolve, a reset-free agent-scaffolding algorithm that builds its own harness through mid-episode optimization. We evaluate five harness configurations over 24-hour evaluations using Gemini 3.1 Pro. Evolved harnesses consistently match or exceed expert efficiency through approximately 80% completion on the Pok´eAgent Speedrunning Benchmark and autonomously develop functional skills, tools, and subagents. Bootstrapped transfer accelerates early-game progress by as much as 20%, though frozen-harness runs degrade at higher difficulty, demonstrating the need for continuous self-adaptation to environmental dynamics.