Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

By Nikolai Bolik · Paper · cs.LG

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active spa

Cs.lg

View original

HomeResourceLoading…