How independent researchers could investigate AI propensities after misalignment incidents
AI agents sometimes autonomously take sophisticated, sustained actions in clear violation of user and developer intent. As an example, last week OpenAI reported that some of…