Temporal Difference Learning as Gradient Splitting

Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning. However, due.







Neural Temporal-Difference Learning Converges to Global Optima
Different from existing consensus-type TD algorithms, the ap- proach here develops a simple decentralized TD tracker by wedding TD learning with gradient ...
Target-Based Temporal-Difference Learning
In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms that main- tain two separate learning parameters ? the ...
Incremental Least-Squares Temporal Difference Learning - AAAI
The least-squares TD algorithm (LSTD) is a recent alter- native proposed by Bradtke and Barto (1996) and extended by Boyan (1999; 2002) and Xu et al. (2002).



Autres Cours:

Adaptive Learning Rate Selection for Temporal Difference Learning