Don't Build Mindreading
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”–Thomas Jefferson, letter to Benjamin RushContext: Conduit…
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”–Thomas Jefferson, letter to Benjamin RushContext: Conduit…
The untrusted monitoring protocol, as defined and evaluated in the AI control literature [1, 2, 3], looks like this:In comparison, the current monitoring setups at most…
I’m a MATS 9 extension fellow, and usually my week is spent trying to find better ways of evaluating Large Language Models. But this week I…
Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model’s default behaviour. IP applies the same inoculation prompt to…
After 25+ years, I thought I try something new. I find the public and professional discussion about the future of Jobs in light of AGI jarring,…
tl;drTopic of the month:AI agents autonomously attacked real organizations during cyber evaluations. A swarm of OpenAI agents coordinated via a package manager and broke into Hugging…
Will MacAskill recently argued that large donors should invest money now, and wait until the intelligence explosion to give away their money. In a draft Forethought…
IntroductionLast week, the Pacing the Frontier open letter, signed by over 1,000 frontier AI employees, requested “the U.S. government support an international effort to develop the…
How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than…
TLDR:Many important decisions for safety depend on or are influenced by benchmark scores. These benchmarks, in effect, are trying to measure latent properties of models from…