Persistent Latent Misalignment, a new dimension of misalignment?
A new paper was released at ICML that I'm worried will open an entire new dimension of alignment problems:Latent Collaboration in Multi-Agent Systems (LatentMAS)TLDR: they show…
A new paper was released at ICML that I'm worried will open an entire new dimension of alignment problems:Latent Collaboration in Multi-Agent Systems (LatentMAS)TLDR: they show…
Note: the modeling assumptions and conclusion are Thomas Kwa's opinion, and others at METR disagree. [1] Also, the math was checked by Claude but not a…
This is the first entry in a sequence of posts which compare a mathematical theory of attention against trained transformers.Links: [GitHub repository]: The code for these…
Below is a summary of my solution to BlueDot’s TAIS Puzzle #1, you can check my blog for the full write-up. The puzzle gave you a…
TLDR:On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked…
I’ve reached a point in my life where I realize that everything I want sits on the other side of action. Earlier this week, I randomly…
Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex…
Right now, most AI safety talent is concentrated in just a few cities: San Francisco, Berkeley, London, Oxford, and Washington D.C. Some other cities also deserve…
Simple mathematical models are almost always too simple to model complex phenomena in the real world, and the one I want to discuss in this post…
The most popular take on the standard free will debate is that you are the algorithm. Your preferences and reasoning that determine your actions IS free…