Near, far: Patch-ordering enhances vision foundation models' scene understanding
| Authors |
|
|---|---|
| Publication date | 2025 |
| Book title | The Thirteenth International Conference on Learning Representations |
| Book subtitle | ICLR 2025 |
| ISBN (electronic) |
|
| Event | 13th International Conference on Learning Representations, ICLR 2025 |
| Number of pages | 28 |
| Organisations |
|
| Abstract |
We introduce NeCo: Patch Neighbor Consistency, a novel self-supervised training loss that enforces patch-level nearest neighbor consistency across a student and teacher model. Compared to contrastive approaches that only yield binary learning signals, i.e. "attract" and "repel", this approach benefits from the more fine-grained learning signal of sorting spatially dense features relative to reference patches. Our method leverages differentiable sorting applied on top of pretrained representations, such as DINOv2-registers to bootstrap the learning signal and further improve upon them. This dense post-pretraining leads to superior performance across various models and datasets, despite requiring only 19 hours on a single GPU. This method generates high-quality dense feature encoders and establishes several new state-of-the-art results such as +2.3 % and +4.2% for non-parametric in-context semantic segmentation on ADE20k and Pascal VOC, +1.6% and +4.8% for linear segmentation evaluations on COCO-Things and -Stuff and improvements in the 3D understanding of multi-view consistency on SPair-71k, by more than 1.5%.
|
| Document type | Conference contribution |
| Language | English |
| Published at |
https://openreview.net/forum?id=Qro97zWC29
(Final published version)
|
| Other links | |
| Downloads |
6193_Near_far_Patch_ordering_e
(Final published version)
|
| Permalink to this page | |