Publication: Learning Optimal Fencing Strategies using Reinforcement Learning in a Simulated MuJoCo Environment
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
High-level fencing demands a complex integration of balance, locomotion, weapon control, and timing, making it a compelling test case for whether structured athletic strategy can emerge from physical constraints and competitive pressure alone, rather than from explicit motion programming. This thesis develops a physics-based Olympic ´ep´ee fencing simulation using MuJoCo, in which a humanoid agent is trained from scratch using PPO self-play, guided by a nine-term shaped reward function that balances dense shaping feedback with the sparse goal of scoring a touch. After 50 million training steps, the resulting policy reliably scores in roughly 80% of episodes and produces human-recognizable behaviors that were not explicitly rewarded, including a stable en garde stance and a forward weight transfer motion qualitatively resembling a fencing lunge. However, more sophisticated tactics that require modeling the opponent’s intentions, such as feints and parries, did not emerge, a gap attributable to the memoryless policy architecture, the first-touch episode termination rule, and simplifications in the humanoid model. These results support a partial version of the emergence hypothesis: biomechanical constraints and adversarial optimization are sufficient to produce recognizable athletic primitives, but opponent-aware strategic behavior appears to require additional structural elements beyond reactive reward optimization.