Skip to content
LessWrong AI · Communities

Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs

TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM’