Skip to content
arXiv cs.AI · Papers

Human-like Object Grouping in Self-supervised Vision Transformers

arXiv:2603.13994v3 Announce Type: replace-cross Abstract: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introd