ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

By Yao Xiao · Paper · cs.CV

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an

AI Infra · Cs.cv

View original

HomeResourceLoading…