Publication:

Learning Optimal Fencing Strategies using Reinforcement Learning in a Simulated MuJoCo Environment

Loading...
Thumbnail Image

Files

written_final_report.pdf (2 MB)

Date

2026-04-16

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

High-level fencing demands a complex integration of balance, locomotion, weapon control, and timing, making it a compelling test case for whether structured athletic strategy can emerge from physical constraints and competitive pressure alone, rather than from explicit motion programming. This thesis develops a physics-based Olympic ´ep´ee fencing simulation using MuJoCo, in which a humanoid agent is trained from scratch using PPO self-play, guided by a nine-term shaped reward function that balances dense shaping feedback with the sparse goal of scoring a touch. After 50 million training steps, the resulting policy reliably scores in roughly 80% of episodes and produces human-recognizable behaviors that were not explicitly rewarded, including a stable en garde stance and a forward weight transfer motion qualitatively resembling a fencing lunge. However, more sophisticated tactics that require modeling the opponent’s intentions, such as feints and parries, did not emerge, a gap attributable to the memoryless policy architecture, the first-touch episode termination rule, and simplifications in the humanoid model. These results support a partial version of the emergence hypothesis: biomechanical constraints and adversarial optimization are sufficient to produce recognizable athletic primitives, but opponent-aware strategic behavior appears to require additional structural elements beyond reactive reward optimization.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation