Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

Researchers proposed a new algorithm, Regularized Emphatic Temporal-Difference Learning (RETD), to stabilize the expected off-policy TD update in reinforcement learning. RETD is designed to improve stability and convergence in scenarios with constant step sizes.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.