The Next Ecology
When I started writing about AI, my concern was ASI. I'm still concerned about AI, but recent events have made me realize we're potentially facing something…
When I started writing about AI, my concern was ASI. I'm still concerned about AI, but recent events have made me realize we're potentially facing something…
Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem. We did not have that proof. An…
In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct…
TL;DR: We’ll probably be getting more AI safety incidents, so amplify the ones that would justify or highlight the urgency of your preferred policy solutions even…
TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust…
“Diversification internationally," Canadian Prime Minister Mark Carney remarks, “is not just economic prudence — it is the material foundation for honest foreign policy.” Carney’s address to…
A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was…
Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to…
In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from…
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all…
LessWrong AI is one of 177 primary AI sources we aggregate. 856 stories from this source have been indexed. Domain: www.lesswrong.com. All posts here link straight to the original — we don't republish content, we point readers at it.
See the full source catalogue or browse by model.