Skip to content
arXiv cs.CL · Papers

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

arXiv:2607.14111v1 Announce Type: new Abstract: Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept vectors into a model's residual stream and measuring whether the model can accurately repor