The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The…