Skip to content
LessWrong AI · Communities

Where does hint-following and concealment arise? A case study on OLMo-3 checkpoints

This work was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I'm grateful to Harshul Basava and Vanessa Ng for their feedback. This is an extension to a prior replication which can be found here.Introduction and MotivationIn an earlier post, I showed that the “