General capability – and capabilities generally – have no good y-axis
BLUF:To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big…
BLUF:To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big…
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems.…
Working on AI Safety, I spend a significant part of my time thinking about how frontier AI Safety research eventually gets translated into deployable safety systems.…
I, for one, do not look forward to sharing the apparent fate of mathematicians. Contrary to my flippant response this morning on X, though, it is…
Honorary rationalist Ferrett Steinmetz describes how he fell down a rabbit hole watching YouTube videos (on mechanical watches) and ended up being "brainwatched" - absorbing ideas…
Work done at Redwood Research, quick update on results found as part of a larger project. Thanks to @SebastianP and @egan for comments on drafts.TL;DRChanging the…
Generally, when you have an identity of X, you are likely to be influenced to stay within the socially acceptable boundaries of identity X. Identity boundaries…
SIA is true when there are no duplicates to worry about, and though it has issues around duplicate creation, so does every other theory of anthropic…
What’s the biggest thing you think you can take in a fight? According to a YouGov poll, 6% of Americans reckoned they could beat a grizzly…
I want to discuss and brainstorm a counterintuitive approach to AI alignment:Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment.To prevent this…