Towards Robustness against Typographic Attack with Training-free Concept Localization
By Bohan Liu · Paper · cs.CV
Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appea