Abstract
We prove under commonly used assumptions the convergence of episodic, that is, sequence-based, actor-critic-like reinforcement algorithms for which the policy becomes more greedy during learning. The most prominent example of such algorithms is the recently introduced RUDDER method, which speeds up the learning of delayed reward problems by reward redistribution. RUDDER is based on simultaneously learning a policy and a reward redistribution network similar to actor-critic methods. We show the convergence of RUDDER which can be generalized to similar actor-critic-like algorithms. In contrast to previous convergence proofs for actor-critic-like methods, we consider whole episodes as learning examples, undiscounted reward, and a policy that becomes more greedy during learning. We employ recent techniques from two time-scale stochastic approximation theory which are equipped with a controlled Markov process to account for the policy getting more greedy. We expect our framework to be useful to prove convergence of other algorithms based on reward shaping or on attention mechanisms.
| Originalsprache | Englisch |
|---|---|
| Titel | Neural Information Processing Systems Foundation (NeurIPS 2019), 2019 |
| Seitenumfang | 8 |
| Publikationsstatus | Veröffentlicht - 2019 |
Wissenschaftszweige
- 305907 Medizinische Statistik
- 202017 Embedded Systems
- 202036 Sensorik
- 101004 Biomathematik
- 101014 Numerische Mathematik
- 101015 Operations Research
- 101016 Optimierung
- 101017 Spieltheorie
- 101018 Statistik
- 101019 Stochastik
- 101024 Wahrscheinlichkeitstheorie
- 101026 Zeitreihenanalyse
- 101027 Dynamische Systeme
- 101028 Mathematische Modellierung
- 101029 Mathematische Statistik
- 101031 Approximationstheorie
- 102 Informatik
- 102001 Artificial Intelligence
- 102003 Bildverarbeitung
- 102004 Bioinformatik
- 102013 Human-Computer Interaction
- 102018 Künstliche Neuronale Netze
- 102019 Machine Learning
- 102032 Computational Intelligence
- 102033 Data Mining
- 305901 Computerunterstützte Diagnose und Therapie
- 305905 Medizinische Informatik
- 202035 Robotik
- 202037 Signalverarbeitung
- 103029 Statistische Physik
- 106005 Bioinformatik
- 106007 Biostatistik
JKU-Schwerpunkte
- Digital Transformation
Dieses zitieren
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver