the polysemanticity of polysemanticity in language models
Polysemanticity is one of the most important concepts when it comes to mechanistic interpretability, basically, studying the internal representations of neural networks. For some reason, the…
Polysemanticity is one of the most important concepts when it comes to mechanistic interpretability, basically, studying the internal representations of neural networks. For some reason, the…
By Sohybe Ibrahim Abdelwahab Amer | June 2026The Google DeepMind mechanistic interpretability team (Neel Nanda et al.) suggested a deliberate shift; instead of relying on reverse-engineering…
OpenClaw scared me a bit and struck me as an "agentic is here" moment. Had I made predictions on what the last 7+ months would look…
There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here. I…
By Roland Pihlakas and Jan Llenzl DagohoyThis post is a slightly updated copy of our Arxiv preprint available at https://arxiv.org/abs/2605.21401 . The tables are converted to…
A couple weeks ago I wrapped production on the audiobook version of Fundamental Uncertainty. It’s now finally starting to be available through various sellers.Right now you…
Benji Berczi, Kyuhee Kim, James Requeima, Sid Black, Cozmin UdudecThis is work done by Benji and Kyuhee during MATS Winter 2026, mentored by Cozmin Ududec, and…
Currently, alignment evaluation works by constructing a situation, observing the model's behavior and scoring it. We put a lot of thought into designing these benchmarks, and…
TL;DR: Current LLMs are bad communicators relative to their agentic capabilities. I claim that articulacy is useful (and perhaps necessary) for AI safety and suggest a…
IntroductionThe aim of this post is to share a quick attempt at grokking the conceptual ideas that lie behind the notion of J-space and how it…