Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Researchers introduce BATON, a dual-axis policy optimization framework for LLM agents, combining Bayesian Feedback Attribution and Trajectory Mass Normalization. Experiments show that both axes provide independent gains and their combination achieves the strongest performance across model scales.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.