Four LLM loss functions → four flavors of LLM misalignment
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the…
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the…
TL;DRHow can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A…
I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building…
This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order…
TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their…
It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with…
Richard Ngo challenged me to set a time box and write down as many of the most important features of my formal epistemology as I can…
When we get explicit strong generalisation to work (see the first post on the matter and the second) my dream would be to create pre-aligned generalising…
A human superpower hidden from even ourselvesI though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it…
I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human…
Alignment Forum is one of 177 primary AI sources we aggregate. 38 stories from this source have been indexed. Domain: www.alignmentforum.org. All posts here link straight to the original — we don't republish content, we point readers at it.
See the full source catalogue or browse by model.