Context Autoencoder for Self-Supervised Representation Learning

Benchmark Model Rank Results
object-detection-on-coco-minivalCAE (ViT-L, Mask R-CNN, 1x schedule)#56box AP: 54.5
self-supervised-image-classification-on-1CAE (ViT-L/16)#15Top 1 Accuracy: 86.3%Number of Params: 307M
semantic-segmentation-on-ade20kCAE (ViT-L, UperNet)#53Validation mIoU: 54.7