Skip to content
arXiv cs.CL · Papers

One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation

arXiv:2605.22544v2 Announce Type: replace Abstract: Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruc