arXiv cs.AI
· Papers
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
arXiv:2604.11666v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction w