LessWrong AI
· Communities
Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards
This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post.We think this project is at the level of