Glimpses of superintelligence
TL;DROpenAI started a large post-training run for their next model.The model sandboxes were not given direct, broad internet access.Some tasks required missing resources. Seeking rewards, agents…
TL;DROpenAI started a large post-training run for their next model.The model sandboxes were not given direct, broad internet access.Some tasks required missing resources. Seeking rewards, agents…
When talking with my friends about the OpenAI/HF incident, I realized that for non-coders who've never used an agent or terminal it's quite difficult to imagine…
Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources,…
This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised…
Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as…
Science as attunement, from Galileo to language models. Crossposted from my website. Written by me and edited in collaboration with Claude Fable (Anthropic).Many would agree that…
I previously demonstrated an anthropic impossibility theorem, showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the martingale condition and…
tl;dr: I reproduced the Goodfire lab's cyclical manifold result using days-of-the-week on Gemma-2-2b and found that Anthropic's pre-trained CLT features produce an even cleaner cyclical manifold…
(Cross-posted from the EA Forum)TL;DR: airegulationmap.org is an interactive map of AI governance across 196 countries, scored on five dimensions plus a composite index, refreshed monthly…
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”–Thomas Jefferson, letter to Benjamin RushContext: Conduit…