Publication:

BUILDING HYDRA: DYNAMIC REINFORCEMENT LEARNING CONTROL FOR OPTIMIZING H2O2 ELECTROSYNTHESIS

Loading...
Thumbnail Image

Files

ACHUCARRO_THESIS_VF.pdf (10.33 MB)Embargo until 2027-07-01

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Hydrogen peroxide (H2O2) is a widely used oxidant whose current industrial production relies almost exclusively on the energy-intensive and centralized anthraquinone process, a method that presents various safety and sustainability challenges. Electrochemical synthesis via the two-electron oxygen reduction reaction offers a distributed, renewable energy compatible alternative, but its implementation is limited by competing reaction pathways that reduce selectivity and by the impracticality of manually tuning the voltage-time pulse parameters needed to optimize production in real time. This thesis presents HYDRA - Hydrogen Peroxide Dynamic RL Agent - a reinforcement learning inspired controller that autonomously tunes voltage-time parameters for adaptive H2O2 electrosynthesis. Adapted from the AlphaFC actor-critic framework originally developed for methanol fuel cells in Dr. Ju Li’s lab at MIT, HYDRA introduces a sequential database structure that captures the causal relationships of physical electrochemical systems, a current-based reward design, and a staged validation methodology that confirms each system component before the next is actuated. Physical validation experiments confirmed that the CB-10% carbon black catalyst drives meaningful oxygen reduction activity, and that pulsed voltage operation produces 14% higher H2O2 concentrations than constant voltage operation under identical conditions, establishing the basis of the set-up to be optimized. The HYDRACritic network, trained on 44 pseudo-exploration samples with data augmentation, achieved an R² of 0.691 under the selected reward strategy, sufficient for initial deployment. During a 30 minute closed-loop deployment on a physical electrolyzer, the Agent was able to carry out voltage-time pulses that resulted in production of H2O2, following it, the Critic was retrained on the expanded dataset and achieved an R2 of 0.968. This demonstrates HYDRA’s capacity for self-improvement through accumulated real experimental data. Most significantly, H2O2 was produced under autonomous RL control during the first deployment. This result has not been demonstrated before in the electrochemical or autonomous control literature. That a controller trained on only 44 samples, operating in its first closed-loop run on physical hardware, was able to drive meaningful electrochemical H2O2 production speaks to the robustness of the framework and the promise of RL-based control for electrochemical synthesis. The results provide a replicable template for deploying reinforcement learning in other dynamic electrochemical systems.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation