Toward A Public Science of Model Behavior
This is a linkpost for our essay "Toward A Public Science of Model Behavior" on transluce.org. The full text is reproduced below.Today’s AI systems frequently behave…
This is a linkpost for our essay "Toward A Public Science of Model Behavior" on transluce.org. The full text is reproduced below.Today’s AI systems frequently behave…
I completed my law degree at a working-class London university. In my first year, I was 18 years old, and I was often the youngest person…
At the end of May, I attended the SciFM26 conference hosted at UChicago which had many speakers discussing agentic science and the future of AI+Science. This…
TL;DR: NLAs are an interesting idea. They do reconstruct the activations well, they can kinda audit model organisms for hidden objectives. But there are hints that…
This work was done as part of the MATS 8.1 Program.0: TL;DRLLMs learn "values": general considerations (e.g. "playfulness & humor", "mental health sensitivity") that influence their…
Some context for this post: I’ve been working part-time as a consultant for the AI Futures Project over the last year. Most of the work I’ve…
In a world where rogue ASI can form a singleton, could it really widely deploy an agent fleet across the world(let alone onto other stellar bodies)…
TL;DR: Some criticisms aimed at FDT are actually aimed at a self-contradictory "unembedded FDT," and are therefore irrelevant to any refutation of FDT.Functional Decision Theory (FDT)…
Indirect Prompt Injection at present day, is one of the main reasons for agentic failures deployed in personal systems as well as enterprise grade applications/systems. The…
Cross-posted from The Foretellix CTO Blog. Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated…