Apple ML Research
· Cloud & Big Tech
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision language models (VLMs). While RL-tuned VLMs improve on visual reasoning benchmarks, they remain vulnerable to weak visual grounding, hallucination