SFF is very suboptimal
I recently served as a recommender in SFF's annual funding round (grants will be decided and announced in September). I'm deeply grateful for SFF's funders, and…
I recently served as a recommender in SFF's annual funding round (grants will be decided and announced in September). I'm deeply grateful for SFF's funders, and…
In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness…
Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation detailsTL;DR: we should cryptographically verify that sub-agent instances/sessions are…
This is a linkpost for Request for Proposals: Research and Applied Work on Digital Minds.The text below was written in the first person by Zach Freitas-Groff,…
“One thing I like about jazz, kid, is that I don’t know what’s going to happen next.” — attributed to Bix Beiderbecke.I hate lectures. I refuse…
This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.TL;DROur theorems and experiments suggest that DPO…
Tl;dr I've been thinking hard and experimenting with how to best take advantage of AI for my work. I've developed some practices, mostly through trial and…
TLDRLLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals. Improving model functional welfare matters…
Summary results/takeawaysI steer Qwen3-32B and 235B along qualia-related emotion directions (blissful, tormented, terrified, serene, etc.) by adding an emotion vector to the residual stream at varying…
In April, Tyler Cowen linked (without comment) to a paper titled "Why AI can simulate but not instantiate consciousness". I've been paying attention to AI and…