Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?
By Philippe Chlenski · Paper · cs.LG
Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make claims