Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

Researchers proposed ConSPO, a new reinforcement learning method that addresses limitations in existing algorithms like GRPO. ConSPO uses length-normalized sequence log-probabilities as rollout scores and contrasts them against negative distractors, leading to improved performance on reasoning benchmarks.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.