arXiv cs.CL
· Papers
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
arXiv:2605.22544v2 Announce Type: replace Abstract: Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruc