Skip to content
arXiv cs.CL · Papers

Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision-making remains underexplored. In this study, we evaluated the performance of three state-of-the-art LLMs-ChatGPT-4o, Ge