Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Researchers introduce BATON, a dual-axis policy optimization framework for LLM agents, combining Bayesian Feedback Attribution and Trajectory Mass Normalization. Experiments show that both axes provide independent gains and their combination achieves the strongest performance across model scales.
Save an API key to vote.