Skip to content
arXiv cs.CL · Papers

Dissociating the Internal Representations of Sycophancy in LLMs

arXiv:2607.07003v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raising the questio