arXiv cs.LG
· Papers
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
arXiv:2510.17426v3 Announce Type: replace-cross Abstract: The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We show that this trade-off can be navigated e