Inoculation Adapters Improve Upon Inoculation Prompting
This is a link post for the paper preprint: Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors from the Center on Long-Term Risk.Selective…
This is a link post for the paper preprint: Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors from the Center on Long-Term Risk.Selective…
As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. Xi gave an important speech yesterday, so this post opens with…
Anthropic recently published Agentic Misalignment Summer 2026The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some…
TL;DRUniversity groups are among the most reliable producers of AI safety talent, yet dozens of top schools that could sustain a group don't have one. We're…
Adapted from my Substack, Funding Anthropalypse.Short version: for organisations aiming for a share of the coming Anthropic and OpenAI windfall - the $37bn+ that could be…
The grabby aliens argumentAccording to standard models of cosmology, there will be habitable planets in our universe for a very long time, and some of these…
I want to break a lance for Jean Piaget. When I come across his name, it's mostly in the context of criticism, much of which misses…
What a time to be alive! Some people posted about how to do cheaper better faster vision and it ended up in fable. Some people posted…
Self-Other Overlap fine tuning described in Carauleanu et al. (2025) attempts to partially fuse the model’s concept of self to its conceptualization of an outside entity…
I’ve recently been rereading Steven Levy’s “Hackers” with my daughter. Levy describes how Brautigan’s 1967 poem “All watched over by machines of loving grace” was inspiring…