OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation
OPD-Aha is a new multimodal on-policy distillation method that improves AI model reasoning by allowing a teacher to evaluate student trajectories using visual evidence. The method reconstructs the distillation target directly from the teacher's visual preference, suppressing continuations that contradict the image and enabling reflection tokens like 'wait' and 'actually'. This approach leads to broad and consistent improvements across various benchmarks.
Save an API key to vote.