Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
By Hiskias Dingeto · Paper · cs.AI
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the recon