Childhood and Education #20: Phones and Screens
We have a respite, so I thought I’d tackle various thoughts on children, phones and screens. GPT-5.6-Sol drops tomorrow, and the Fable agents are hard at…
We have a respite, so I thought I’d tackle various thoughts on children, phones and screens. GPT-5.6-Sol drops tomorrow, and the Fable agents are hard at…
Summary: Right now, we have no idea how practical it is for countries to sabotage each others’ AI projects. This makes it hard to forecast what…
I wrote this piece while at the AFFINE Superintelligence Alignment Seminar during discussions about the difference between AI Alignment and AI Safety. If you’re simply interested…
Subliminal learning is the phenomenon where a language model picks up a behavioral trait—such as fondness for cats—by training on data from a trait-carrying teacher that…
tl;drPaper of the month:Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition —…
This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from…
This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘proto-training gaming,’…
For something that “simply predicts the next token”, Large Language Models are a surprisingly versatile technology. In addition to the obvious applications, such as chatbots, translation…
Thanks to @eigengender and especially Chris Lakin and Simon Dima for valuable comments on a draft.In this post, I will present an alternate framing[1] of LessWrong-style…
Figure from Chapter 4 of Understanding Knowledge by Michael HuemerOn a high level there aren’t that many ways knowledge can be justified (or not). This figure…