EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG-Tree is a new algorithm for gradient-efficient reinforcement learning, specifically designed for language models. It improves policy-gradient estimation by allocating branches in a tree-based rollout construction to reduce uncertainty about the policy gradient per unit of compute. This can lead to better gradient estimation and improved performance in tasks such as Wordle and cloned-state control.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.