3D-Aware VLMs with Implicit and Explicit Geometries

By Wenhao Li · Paper · cs.CV

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances

Cs.cv

View original

HomeResourceLoading…