Skip to content
arXiv cs.LG · Papers

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

arXiv:2608.07827v1 Announce Type: new Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical accuracy. Prior work has shown that token probabilities are often overconfident, we investig