OPD-V: Visual On-Policy Self-Distillation with Modality Balance
By Aniri · Paper · cs.CV
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlo