Skip to content
r/MachineLearning · Communities

Mechanistic interpretability: a first paper on disentangling a convolutional neuron [R]

I have recently started working in mechanistic interpretability independently, starting with distill circuits thread My work is on disentangling and closely studying a single neuron, a 1x1 convolution in inceptionv1 model (and applying the method to other neurons in the same layer). The key insight was that the hadamar