arXiv cs.CV
· Papers
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
arXiv:2607.00125v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains challenging. To bridge this gap, we present DeCoDe, a simple yet effective technique that ena