arXiv cs.CL
· Papers
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference
arXiv:2601.13300v2 Announce Type: replace Abstract: Benchmarking large language models (LLMs) is critical for understanding their capabilities, limitations, and robustness. In addition to interface artifacts, prior studies have shown that LLM decisions can be influenced by directive signals such as social cues, framing