Test-Time Training for Modality Order Consistency in Vision-Language Models

By Aditi Gupta · Paper · cs.CV

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a

Model Launch · Cs.cv

View original

HomeResourceLoading…