DOPD: Dual On-policy Distillation
By Xinlei Yu · Paper · cs.AI
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse
Opd · Distillation · Cs.ai