R-MAE: Regions Meet Masked Autoencoders
| Authors |
|
|---|---|
| Publication date | 2024 |
| Book title | The Twelfth International Conference on Learning Representations |
| Book subtitle | ICLR 2024 |
| ISBN (electronic) |
|
| Event | 12th International Conference on Learning Representations, ICLR 2024 |
| Number of pages | 18 |
| Organisations |
|
| Abstract |
In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.
|
| Document type | Conference contribution |
| Language | English |
| Published at |
https://openreview.net/forum?id=ba84RDHFnz
(Final published version)
|
| Other links | |
| Downloads |
2410_R_MAE_Regions_Meet_Masked
(Final published version)
|
| Permalink to this page | |
