Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short…
TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short…
Working with John Wentworth is confusing and overwhelming at times.[1] The guy has a lot of information and models in his head and he doesn’t always…
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far. Key…
0. Intro Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty…
Read this short exchange. A: "Green apples are delicious." B: "Huh? Aren't they better when they're ripe?" A: "No, I meant Granny Smiths." A said "Green…
From the Mythos preview system card (emphasis mine):We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of…
Blogs have shaped our philosophical worldviews, found us careers and friends, and changed our lives. There’s a good chance that a great blog of yore is…
The standard ML pipeline is well known: you want to teach a task to an algorithm. Take your data which encodes that task, split it into…
TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be…
This week, we're talking about anthropic reasoning. What are we to make of fine-tuned coincidences in science and mathematics? Are they coincidences, signs of structure we…