Skip to content
LessWrong AI · Communities

Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?

This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-relevant indicators change with quantization; either in valence (do indicators shift toward more negative / more distressed / more boundary-eroded states?) or in stability (do indicators become noisier,