Skip to content
arXiv cs.CL · Papers

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often exploit adversarial context to deceive LLMs, prompting them to generate responses