Skip to content
arXiv stat.ML · Papers

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

arXiv:2607.22951v1 Announce Type: new Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estima