Skip to content
r/MachineLearning · Communities

DINOv2 way worse than SigLIP in k-NN. Is this expected? [R]

Doing a bachelor thesis on fine-grained car classification (telling apart VW Golf generations from listing photos). Simple setup: frozen encoder → embeddings → weighted k-NN. On my small dataset (175 train / 132 test): I thought maybe it was a cosine vs euclidean thing, but my embeddings are L2-normalized so both give