Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

By Tadanobu Chuyo Kamijo · Paper · cs.CL

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails

Cs.cl

View original

HomeResourceLoading…