Multimodal Model Diffing for Feature Discovery and Control

By Hunar Batra · Paper · cs.CV

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable featur

Cs.cv

View original

HomeResourceLoading…