Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

A new method, ActObs, is introduced to supervise both agent action tokens and observation tokens during reinforcement learning. This approach improves performance in tasks such as code editing and retains more entropy in the policy, leaving it closer to its initialization.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.