Skip to content
LessWrong AI · Communities

Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information

TLDR: We train a “matryoshka” NLA that, unlike standard NLAs, is trained to put the most important details at the start; it is trained by randomly truncating the verbalizer’s explanations before showing them to the reconstructor. We find that our matryoshka NLA frontloads claims that are important to reconstruction (mo