r/LocalLLaMA
· Communities
Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params
Weights, all 4 sizes, Apache-2.0 (ViT-S / ViT-B / ViT-L / ViT-g): https://huggingface.co/collections/robbyant/lingbot-vision Code: https://github.com/robbyant/lingbot-vision Project page: https://technology.robbyant.com/lingbot-vision Self-supervised DINO-family backbone, but the masking is boundary-driven: the teacher