Fine-Tuning, The Hierarchy Problem, and What Neutrons Tell Us About God
This week, we're talking about anthropic reasoning. What are we to make of fine-tuned coincidences in science and mathematics? Are they coincidences, signs of structure we…
This week, we're talking about anthropic reasoning. What are we to make of fine-tuned coincidences in science and mathematics? Are they coincidences, signs of structure we…
The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240T params in 2028 [1]…
If you're training any type of toy model of superposition, Mean Squared Error (MSE) loss is unusually bad. [1] Related work We aren't the first to…
In a recent post, we presented PIRAMID, its leadership and research pillars, and a plan for how they fit together. In this post, we sketch a…
This is my submission to BlueDot's Technical AI Safety Puzzle #1, for which I received an Honorable Mention. Congratulations to Gustavo Korzune Gurgel, Patryk Perduta (his…
I fucking love em dashes. I spam them all the time. They make me feel free.I don’t give a shit what the “oh yeah let’s examine…
I started writing online daily more than 630 days ago. These are the things I wish I knew at the beginning. 1. You can start with…
I'm uncertain of what to do.Something clean and clear shines out: if people don't see any more of my slavery posts, will they think that slavery…
TLDR -We talk about scheming, and why research on this phenomenon is crucial for AI safety.We find a particular environment/scenarion where scheming happens at a higher…
(Adapted from a post on my Substack.)A recent Washington Post tech report “Are ChatGPT and other AI chatbots politically biased? We tested them” went viral with…