Measuring Reward-Seeking by Instilling Contrastive Beliefs
This is an unofficial automated linkpost. Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that,…
This is an unofficial automated linkpost. Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that,…
We recently published our paper on "Measuring Reward-Seeking via Contrastive Belief Updates". We're excited about research like this, and there are many more open problems than…
There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can…
Money is pouring in, people are looking for new areas to fund, and the invisible hand is starting to grab a bit at AI for epistemics…
A mini-series, continuing my efforts to document the AI landscape as of July 2026Part 1 - The Dangers of LLMs (you are here)Part 2 - The…
In July 2022 I was in a parking lot with a Portuguese colleague, trying to fix the cargo-metering system of a 12-ton tanker truck. During a…
Wait so is it true that next week the largest open weight model will be 2-3x bigger than last week? I don't like that.. it was…
I came across this study on X and tested it myself; the findings align with the results: FSRS is more recognizable than Jarrett Ye among LLMs.…
BackgroundThe AI safety field has spent a decade building tools for systems trained to think and digitally act. The next decade will likely deliver widely deployed…
In this post, I propose adapting banking risk management frameworks (specifically capital adequacy requirements like Basel III) to frontier AI labs. By forcing them to hold…