This thesis explores how Reinforcement Learning (RL), and in particular how deep neural network architectures can be used to learn optimal trading and optimal execution strategies in a complex environment, such as the financial market. Moreover, we show how these learning algorithms can shed light on how multiple financial agents interact when market impact is generated by their trading. Therefore, the first result presented in this thesis, shows that Double Deep Q-Learning (DDQL) can successfully recover optimal trading strategies for execution problems when market liquidity is latent and time-varying, both in deterministic and stochastic settings. In fact, considering an Almgren-Chriss environment with temporary and permanent impacts evolving according to different dynamics, we show that DDQL learns the analytical optimal solution whenever it exists, and outperforms standard benchmarks and approximations otherwise. These findings highlight the ability of an agent, modelled with deep RL, to infer optimal behaviour when the underlying market dynamics are unknown, moreover help shed lights on how the behaviour of the agent changes, in terms of performance, when the information provided changes. In the second contribution, we extend the use of RL to trading environments driven by latent information. We study an optimal trading problem where the trading signal follows an Ornstein--Uhlenbeck process with Markovian regime-switching dynamics for its parameters. To extract information from these latent regimes, and thus find the optimal trading strategy, we combine the Deep Deterministic Policy Gradient (DDPG) algorithm with Recurrent Neural Networks (RNNs), specifically Gated Recurrent Units (GRUs), to model temporal dependencies and the hidden structure that drives the evolution of the trading signal. We design, develop and compare the performance in terms of trading results of three novel architectures: a one-step approach that directly encodes GRU hidden states (hid-DDPG), and two two-step approaches that respectively rely on posterior regime probabilities (prob-DDPG) or forecasts of the next signal value (reg-DDPG). Through extensive numerical experiments and an empirical application to equity pair trading, we find that probabilistic inference of hidden regimes significantly enhances the performance, along with the interpretability, of RL-based trading strategies. Finally, the third contribution of the thesis examines how autonomous traders, modelled with deep RL, interact in the same market where market impact is present. By modelling two agents with DDQL in a multi-agent Almgren-Chriss framework, we investigate how independent learning dynamics may deviate from traditional game-theoretic equilibria, resulting in supra-competitive equilibria. In fact, our results based on extensive numerical simulations show that, instead of converging to the Nash equilibrium, agents learn on average supra-competitive strategies, aligning closely with the Pareto-optimal solution to the impact game considered. Furthermore, we explore how changes in market volatility affect both the stability and convergence of learned equilibria. Overall, this thesis studies how deep RL can be a powerful tool for optimal execution, market microstructure, and to study the emergent behaviour of autonomous financial agents. Therefore, using tools from stochastic control, econometrics, and neural network design, we provide novel insights into how RL can learn to act optimally under uncertainty, adapt to latent dynamics, and interact strategically in complex market environments.

Reinforcement learning for optimal execution, trading and impact games / Macrì, Andrea; relatore: LILLO, FABRIZIO; Scuola Normale Superiore, ciclo 37, 19-Jun-2026.

Reinforcement learning for optimal execution, trading and impact games

MACRÌ, Andrea
2026

Abstract

This thesis explores how Reinforcement Learning (RL), and in particular how deep neural network architectures can be used to learn optimal trading and optimal execution strategies in a complex environment, such as the financial market. Moreover, we show how these learning algorithms can shed light on how multiple financial agents interact when market impact is generated by their trading. Therefore, the first result presented in this thesis, shows that Double Deep Q-Learning (DDQL) can successfully recover optimal trading strategies for execution problems when market liquidity is latent and time-varying, both in deterministic and stochastic settings. In fact, considering an Almgren-Chriss environment with temporary and permanent impacts evolving according to different dynamics, we show that DDQL learns the analytical optimal solution whenever it exists, and outperforms standard benchmarks and approximations otherwise. These findings highlight the ability of an agent, modelled with deep RL, to infer optimal behaviour when the underlying market dynamics are unknown, moreover help shed lights on how the behaviour of the agent changes, in terms of performance, when the information provided changes. In the second contribution, we extend the use of RL to trading environments driven by latent information. We study an optimal trading problem where the trading signal follows an Ornstein--Uhlenbeck process with Markovian regime-switching dynamics for its parameters. To extract information from these latent regimes, and thus find the optimal trading strategy, we combine the Deep Deterministic Policy Gradient (DDPG) algorithm with Recurrent Neural Networks (RNNs), specifically Gated Recurrent Units (GRUs), to model temporal dependencies and the hidden structure that drives the evolution of the trading signal. We design, develop and compare the performance in terms of trading results of three novel architectures: a one-step approach that directly encodes GRU hidden states (hid-DDPG), and two two-step approaches that respectively rely on posterior regime probabilities (prob-DDPG) or forecasts of the next signal value (reg-DDPG). Through extensive numerical experiments and an empirical application to equity pair trading, we find that probabilistic inference of hidden regimes significantly enhances the performance, along with the interpretability, of RL-based trading strategies. Finally, the third contribution of the thesis examines how autonomous traders, modelled with deep RL, interact in the same market where market impact is present. By modelling two agents with DDQL in a multi-agent Almgren-Chriss framework, we investigate how independent learning dynamics may deviate from traditional game-theoretic equilibria, resulting in supra-competitive equilibria. In fact, our results based on extensive numerical simulations show that, instead of converging to the Nash equilibrium, agents learn on average supra-competitive strategies, aligning closely with the Pareto-optimal solution to the impact game considered. Furthermore, we explore how changes in market volatility affect both the stability and convergence of learned equilibria. Overall, this thesis studies how deep RL can be a powerful tool for optimal execution, market microstructure, and to study the emergent behaviour of autonomous financial agents. Therefore, using tools from stochastic control, econometrics, and neural network design, we provide novel insights into how RL can learn to act optimally under uncertainty, adapt to latent dynamics, and interact strategically in complex market environments.
19-giu-2026
Settore SECS-S/06 - Metodi mat. dell'economia e Scienze Attuariali e Finanziarie
Metodi computazionali e modelli matematici per le scienze e la finanza
37
Deep Reinforcement Learning; Market Microstructure; Market Impact Games; Optimal Execution; Optimal Trading; Machine Learning
LILLO, FABRIZIO
Scuola Normale Superiore
File in questo prodotto:
File Dimensione Formato  
Macri_tesi_DEF.pdf

accesso aperto

Descrizione: Tesi
Tipologia: Published version
Licenza: Non specificata
Dimensione 7.6 MB
Formato Adobe PDF
7.6 MB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11384/171884
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact