Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers
Researchers propose a method to improve offline on-policy distillation for AI models, allowing them to learn from imperfect teacher supervision and reducing the need for additional generation. This can improve performance and efficiency in tasks like code generation and mathematical reasoning.
Save an API key to vote.