Skip to content
LessWrong AI · Communities

We're talking past our models; or, How a model defined its "evil" vector as dread

SummaryWe train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Respons