Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
By Shayan Talaei · Paper · cs.CL
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale. Such preferential biases can be introduced by any actor in the model's supply chain and are most dangerous when the model reveals its pre