Skip to content
arXiv cs.CL · Papers

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

arXiv:2608.09900v2 Announce Type: replace Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safet