Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Benchmark Model Rank Results
image-classification-on-imagenetdata2vec 2.0#92Top 1 Accuracy: 87.4%