arXiv stat.ML
· Papers
Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model
arXiv:2607.22951v1 Announce Type: new Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estima