Fixing rewards for NLA to reduce confabulation
Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full…
Hello,This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place.Note: 1. This post is 100% human-written. 2. Full…
Epistemic status: this is research engineering, not mechanistic interpretability . The compute/cost claims are measured or derived from architecture constants. The quality claims (faithfulness comparisons, spectral…
Over three years ago, I first considered pulling my personal fire alarm. I think I'm now ready to do it. What is my reasoning? In the…
I'm trying to get better at thinking, communicating and arguing. So I've started a substack. I'm looking for feedback/engagement/advice/criticism, if anyone is willing to provide.==========George Hotz…
tl;dr: some meditations on the shift away from probabalistic/frequentist reasoning to vibes-based means for quantifying and understanding big risks. The vibes in an email can be…
TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically…
Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set…
Authors: Theresa G., Simon S., Siva Kumar Lakkoju.Epistemic status/effort: exploratory red-teaming as part of a two-day hackathon during ARENA 8.0. Our attacks can be reproduced based…
We return to Bold Monk brewing for a vigorous discussion of rationalism and whatever else we deem fit for discussion – hopefully including actual discussions of…
This is a research summary for an ongoing project I am working on as part of the UChicago Existential Risks Laboratory Summer Research Fellowship. I would…