Model Hypnosis: Strong control of AI via additive subliminal effects

By Enric Boix-Adsera · Paper · cs.CL

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and

Cs.cl

View original

HomeResourceLoading…