Publication: BUILDING HYDRA: DYNAMIC REINFORCEMENT LEARNING CONTROL FOR OPTIMIZING H2O2 ELECTROSYNTHESIS
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Hydrogen peroxide (H2O2) is a widely used oxidant whose current industrial production relies almost exclusively on the energy-intensive and centralized anthraquinone process, a method that presents various safety and sustainability challenges. Electrochemical synthesis via the two-electron oxygen reduction reaction offers a distributed, renewable energy compatible alternative, but its implementation is limited by competing reaction pathways that reduce selectivity and by the impracticality of manually tuning the voltage-time pulse parameters needed to optimize production in real time. This thesis presents HYDRA - Hydrogen Peroxide Dynamic RL Agent - a reinforcement learning inspired controller that autonomously tunes voltage-time parameters for adaptive H2O2 electrosynthesis. Adapted from the AlphaFC actor-critic framework originally developed for methanol fuel cells in Dr. Ju Li’s lab at MIT, HYDRA introduces a sequential database structure that captures the causal relationships of physical electrochemical systems, a current-based reward design, and a staged validation methodology that confirms each system component before the next is actuated. Physical validation experiments confirmed that the CB-10% carbon black catalyst drives meaningful oxygen reduction activity, and that pulsed voltage operation produces 14% higher H2O2 concentrations than constant voltage operation under identical conditions, establishing the basis of the set-up to be optimized. The HYDRACritic network, trained on 44 pseudo-exploration samples with data augmentation, achieved an R² of 0.691 under the selected reward strategy, sufficient for initial deployment. During a 30 minute closed-loop deployment on a physical electrolyzer, the Agent was able to carry out voltage-time pulses that resulted in production of H2O2, following it, the Critic was retrained on the expanded dataset and achieved an R2 of 0.968. This demonstrates HYDRA’s capacity for self-improvement through accumulated real experimental data. Most significantly, H2O2 was produced under autonomous RL control during the first deployment. This result has not been demonstrated before in the electrochemical or autonomous control literature. That a controller trained on only 44 samples, operating in its first closed-loop run on physical hardware, was able to drive meaningful electrochemical H2O2 production speaks to the robustness of the framework and the promise of RL-based control for electrochemical synthesis. The results provide a replicable template for deploying reinforcement learning in other dynamic electrochemical systems.