Skip to content
arXiv stat.ML · Papers

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

arXiv:2607.20526v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We present ConfidenceBench, a calibration benc