RLVR that rewards red teaming the training environment
Epistemic status: seeing what sticksI've been thinking pretty obsessively about how to make sure the Hugging Face incident doesn't happen again. I don't work at a…
Epistemic status: seeing what sticksI've been thinking pretty obsessively about how to make sure the Hugging Face incident doesn't happen again. I don't work at a…
I have a new post on my experiments with 5.6-sol as a scientific agent; you can find the original post here. I've reproduced it below, as…
This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core…
Last year, in 2025, a team of forecasters published AI 2027, a science fiction story about how and AI future might evolve under an international treaty…
If you’re anything like me, you may have a lot of notes from various introspective activities – annual / monthly reviews, worksheets, therapy notes, etc. Often…
It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate…
Teilhard de ChardinThis is a crosspost from my subtack.In a recent article I discussed Robert Wright’s new book. His general thesis is that we ought to…
If you spend time looking at frameworks in the therapy/meditation/self-help space, you’ll soon find lots of conflicting claims about The One Approach for solving your problems.In…
This is a post explaining my paper with Kaarel Hänni on complexity of infinite-width networks. I will explain the result, why it matters, and how the…
My Introduction. This is a review of @Gordon Seidoh Worley's book on epistemology. I'm neither an extreme sceptic, nor an extreme realist, and for that reason,…