LessWrong AI
· Communities
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
TLDR: We train a “matryoshka” NLA that, unlike standard NLAs, is trained to put the most important details at the start; it is trained by randomly truncating the verbalizer’s explanations before showing them to the reconstructor. We find that our matryoshka NLA frontloads claims that are important to reconstruction (mo