OPD
TeacherOn-policy distillation asks a stronger or frozen teacher to score the student's sampled trajectory with token-level probabilities.
- Best when a reliable teacher is available.
- Useful for recovery distillation after domain updates.
- Can preserve broad behavior while adding new skills.