arXiv cs.CL
· Papers
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more