arXiv cs.LG
· Papers
Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
arXiv:2607.08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured