arXiv cs.CL
· Papers
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
arXiv:2607.14111v1 Announce Type: new Abstract: Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept vectors into a model's residual stream and measuring whether the model can accurately repor