arXiv cs.CL
· Papers
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
arXiv:2602.06854v2 Announce Type: replace Abstract: Multi-turn jailbreaks capture the real threat model for safety-aligned chatbots, where single-turn attacks are merely a special case. Yet existing approaches break under exploration complexity and intent drift. We propose SEMA, a simple yet effective framework that tr