MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
By Changhao Xiang · Paper · cs.CV
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle