arXiv cs.CL
· Papers
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
arXiv:2608.09900v2 Announce Type: replace Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safet