Fine-Grained Multi Image Object Hallucination Benchmark

By Joonki Min · Paper · cs.CV

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions

Model Launch · Cs.cv

View original

HomeResourceLoading…