Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER

Publikation: Preprints, Working Paper und ForschungsberichteVorabpublikation

Abstract

We prove under commonly used assumptions the convergence of actor-critic reinforcement learning algorithms, which simultaneously learn a policy function, the actor, and a value function, the critic. Both functions can be deep neural networks of arbitrary complexity. Our framework allows showing convergence of the well known Proximal Policy Optimization (PPO) and of the recently introduced RUDDER. For the convergence proof we employ recently introduced techniques from the two time-scale stochastic approximation theory. Our results are valid for actor-critic methods that use episodic samples and that have a policy that becomes more greedy during learning. Previous convergence proofs assume linear function approximation, cannot treat episodic examples, or do not consider that policies become greedy. The latter is relevant since optimal policies are typically deterministic.
OriginalspracheEnglisch
Seitenumfang20
DOIs
PublikationsstatusVeröffentlicht - 2020

Publikationsreihe

NamearXiv.org
ISSN (Druck)2331-8422

Wissenschaftszweige

  • 305907 Medizinische Statistik
  • 202017 Embedded Systems
  • 202036 Sensorik
  • 101004 Biomathematik
  • 101014 Numerische Mathematik
  • 101015 Operations Research
  • 101016 Optimierung
  • 101017 Spieltheorie
  • 101018 Statistik
  • 101019 Stochastik
  • 101024 Wahrscheinlichkeitstheorie
  • 101026 Zeitreihenanalyse
  • 101027 Dynamische Systeme
  • 101028 Mathematische Modellierung
  • 101029 Mathematische Statistik
  • 101031 Approximationstheorie
  • 102 Informatik
  • 102001 Artificial Intelligence
  • 102003 Bildverarbeitung
  • 102004 Bioinformatik
  • 102013 Human-Computer Interaction
  • 102018 Künstliche Neuronale Netze
  • 102019 Machine Learning
  • 102032 Computational Intelligence
  • 102033 Data Mining
  • 305901 Computerunterstützte Diagnose und Therapie
  • 305905 Medizinische Informatik
  • 202035 Robotik
  • 202037 Signalverarbeitung
  • 103029 Statistische Physik
  • 106005 Bioinformatik
  • 106007 Biostatistik

JKU-Schwerpunkte

  • Digital Transformation

Dieses zitieren