MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

By Changhao Xiang · Paper · cs.CV

Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle

Cs.cv

View original

HomeResourceLoading…