Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
By Raphaël Bonnet-Guerrini · Paper · astro-ph.HE
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the mode