Entanglement Between an AI and Its Environment
TL;DR:We introduce a concept we call "entanglement": roughly, the amount of information that the AI has about its environment.We distinguish between "actual" entanglement between a specific…
TL;DR:We introduce a concept we call "entanglement": roughly, the amount of information that the AI has about its environment.We distinguish between "actual" entanglement between a specific…
I just had one of those delightful moments where I have a very specific idea, and then I search for it (Claude Research in this case),…
We believe that ensuring AI goes well does not only require technical skills. Every technical choice rests on philosophical assumptions that technical training does not prepare…
SummaryLLMs have internal emotion representations that causally shape their behaviour. This was recently observed in Claude: positive emotions like happy and loving are linked to sycophancy,…
Slack by Zvi is a much-loved post which captures something important. But the core idea’s contours can be hard to pin down, and people attempting to…
Consider the following dialogue, between two forecasters aiming to predict the next presidential election:Forecaster Alice: I believe that candidate X will win the election, with probability…
A summary of our ICML 2026 paper, Architecture Matters for Multi-Agent Security (Hagag, Anderson, Schroeder de Witt, Scheffler). The scenarios and experiments were built on Orbit,…
Imagine an astronomer who discovers an asteroid with a 50% chance of hitting Earth in 2035. She goes on TV. She testifies before Congress. She founds…
(parts 2 and 3 to follow)Summary of this postThis post is on the results of a mechanistic interpretability project aimed at understanding the internals of Maia…
In keeping with my tradition and given I will lose access to Fabel in a couple days, I have asked fable to read all my organically…