Skip to content
arXiv cs.AI · Papers

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

arXiv:2510.22170v3 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies using situational judgment t