Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos
Open paper
Benchmark
Model
Rank
Results
long-video-retrieval-background-removed-on
MCN
#4
Cap. Avg. R@1: 53.4
Cap. Avg. R@5: 75.0
Cap. Avg. R@10: 81.4
Rank counts only results with a code link.