Test-Time Training for Modality Order Consistency in Vision-Language Models
By Aditi Gupta · Paper · cs.CV
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a