Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
Researchers proposed a new algorithm, Regularized Emphatic Temporal-Difference Learning (RETD), to stabilize the expected off-policy TD update in reinforcement learning. RETD is designed to improve stability and convergence in scenarios with constant step sizes.
Save an API key to vote.