VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]
While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive…