Alignment Forum
· Communities
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values
TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities wh