Skip to content
arXiv cs.AI · Papers

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

arXiv:2604.11666v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction w