A Dynamical Model of AI Governability
A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence…
A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence…
Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training
Interim report on ongoing work on reward hacking
Announcing Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
Research update on on applying local volume measurement to downstream tasks
In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas…
Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Using Product Key Memories to encode sparse coder features
In this post, we show that when two TopK SAEs are trained on the same data, with the same batch order but with different random initializations,…
Using interpretations of SAE latents to simulate activations.
EleutherAI is one of 177 primary AI sources we aggregate. 24 stories from this source have been indexed. Domain: blog.eleuther.ai. All posts here link straight to the original — we don't republish content, we point readers at it.
See the full source catalogue or browse by model.