Skip to content
r/MachineLearning · Communities

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics. Also, clinically meaningful but rare words