Skip to content
LessWrong AI · Communities

Does Your LLM Trust You?

This is a very late post about a project that I did a few months ago as part of the application to Neel Nanda's MATS 10.0 stream. I'm posting the results rather than making a strong claim about any mechanism.Executive SummaryChen et al.[1] demonstrated that LLMs form internal profiles of users from limited context to e