Publication:

Hybrid Reinforcement Learning for Optimal Trade Execution under Transient Market Impact

Loading...
Thumbnail Image

Files

Senior_Thesis_Alexander_Koubaa.pdf (4.8 MB)

Date

2026-04-09

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

This thesis develops a calibrate-simulate-learn pipeline for optimal trade execution under transient market impact. A state-dependent propagator model, in which impact persistence and decay rate are explicit functions of spread, depth, order-flow imbalance, and the Hawkes branching ratio of the arrival process, is calibrated from LOBSTER limit order book data for three liquid U.S. equities subject to no-dynamic-arbitrage constraints. Behavioral cloning and Conservative Q-Learning policies are trained offline on logged trajectories from classical schedule baselines and evaluated via Monte Carlo within the calibrated environment. The behavioral cloning policy reduces mean implementation shortfall by 86 percent relative to TWAP (21.24 versus 148.65 basis points, t = 3.03) and IS variance by 99.2 percent. A three-component IS decomposition identifies the mechanism: the learned policy pays higher instantaneous slippage but eliminates the residual transient impact that TWAP accumulates by executing into its own prior price depression. Feature importance confirms that remaining inventory and time-to-go account for 71.5 percent of predictive weight, with spread and regime label contributing less than 1.5 percent combined. The improvement derives from liquidation geometry in inventory-time space, not microstructure-reactive timing, and is robust to stress perturbations, cross-asset transfer, and calibration sensitivity analysis. All results hold within the calibrated simulator following the principled simulation methodology of Kolm and Westray.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation