Reproducible scaling laws for contrastive language-image learning

Benchmark Model Rank Results
image-classification-on-imagenetOpenCLIP ViT-H/14#43Top 1 Accuracy: 88.5%
open-vocabulary-attribute-detection-on-ovad-1Open CLIP ViT-B32#6mean average precision: 17.0
zero-shot-cross-modal-retrieval-on-flickr30kOpenCLIP VIT-H/14#20Image-to-text R@5: 99.3Text-to-image R@5: 94.1