ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

By Sahil Kale · Paper · cs.CL

Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of indepe

Model Launch · Cs.cl

View original

HomeResourceLoading…