arXiv stat.ML
· Papers
Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
arXiv:2604.10727v2 Announce Type: replace Abstract: Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gaussian losses but can be ineffective when losses or rewards are heavy-tailed. We develop