Multimodal Model Diffing for Feature Discovery and Control
By Hunar Batra · Paper · cs.CV
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable featur