Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
Researchers propose Coupled Calibration and Learning (CCL), an algorithm for LLM distillation that mitigates teacher bias without target-domain reward feedback. CCL calibrates the teacher using source feedback and trains the student on target questions, achieving polynomial convergence to the oracle student. This advancement improves LLM distillation, enabling better transfer of capabilities without transferring biases.
Save an API key to vote.