Reinforcement Learning (RL) provides a standard framework for sequential decision-making, but state-of-the-art Deep RL (DRL) methods are often sample-inefficient and struggle to generalize beyond small-scale training scenarios. We propose a neuro-symbolic DRL approach that integrates background symbolic knowledge to improve sample efficiency and generalization to more complex, unseen tasks. Partial policies learned in simple domains are transferred as logical rules and used for online reasoning to guide learning by biasing exploration and rescaling Q-values during exploitation. This integration enhances interpretability and accelerates convergence, particularly in sparse-reward and long-horizon settings. Experiments show superior performance over state-of-the-art reward machine methods.

Sample-Efficient Neurosymbolic Deep Reinforcement Learning

Veronese, Celeste;Farinelli, Alessandro;Meli, Daniele
2026-01-01

Abstract

Reinforcement Learning (RL) provides a standard framework for sequential decision-making, but state-of-the-art Deep RL (DRL) methods are often sample-inefficient and struggle to generalize beyond small-scale training scenarios. We propose a neuro-symbolic DRL approach that integrates background symbolic knowledge to improve sample efficiency and generalization to more complex, unseen tasks. Partial policies learned in simple domains are transferred as logical rules and used for online reasoning to guide learning by biasing exploration and rescaling Q-values during exploitation. This integration enhances interpretability and accelerates convergence, particularly in sparse-reward and long-horizon settings. Experiments show superior performance over state-of-the-art reward machine methods.
2026
Knowledge Transfer
Neurosymbolic RL
Sample efficiency
File in questo prodotto:
File Dimensione Formato  
RPUM9981.pdf

accesso aperto

Tipologia: Versione dell'editore
Licenza: Creative commons
Dimensione 3.09 MB
Formato Adobe PDF
3.09 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11562/1198607
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact