Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

By Hiskias Dingeto · Paper · cs.AI

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the recon

Cs.ai

View original

HomeResourceLoading…