$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
By Jiawei Xu · Paper · cs.LG
On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisel