OPD-V: Visual On-Policy Self-Distillation with Modality Balance

By Aniri · Paper · cs.CV

On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlo

Cs.cv

View original

HomeResourceLoading…