Skip to content
LessWrong AI · Communities

A Structural Similarity Between Two Open Corrigibility Questions

A corrigible agent "acts opposite the trope of 'be careful what you wish for' by cautiously reflecting on itself as a flawed tool and focusing on empowering the principal to fix its flaws and mistakes." Intuitively, corrigibility is fairly easy to grasp. A corrigible agent should get constant feedback from the user, cl