| Benchmark | Model | Rank | Results |
|---|---|---|---|
| image-classification-on-imagenet | OpenCLIP ViT-H/14 | #43 | Top 1 Accuracy: 88.5% |
| open-vocabulary-attribute-detection-on-ovad-1 | Open CLIP ViT-B32 | #6 | mean average precision: 17.0 |
| zero-shot-cross-modal-retrieval-on-flickr30k | OpenCLIP VIT-H/14 | #20 | Image-to-text R@5: 99.3Text-to-image R@5: 94.1 |