GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

GVPO++ is a novel post-training method for large language models (LLMs) that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme, offering a unique optimal solution and flexible sampling distributions without relying on importance sampling. This method also extends to on-policy distillation and enables the optimization of a broad family of objectives.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.