LessWrong AI
· Communities
"Correct Answer Features" Cannot Explain Multiple Choice Capabilities
TL;DR: I present theoretical and empirical evidence that LLMs cannot be (exclusively) using a "correct answer feature" as the main mechanism by which they perform multiple choice question answering. A hypothetical correct-answer feature would indicate the "correctness" of an option on the final token(s) of that option.