Skip to content
arXiv cs.AI · Papers

Divergent Response Modes in Frontier Language Models Under Steering Pressure

arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six fro