End-to-End Learning of Visual Representations from Uncurated Instructional Videos

Benchmark Model Rank Results
action-recognition-on-rareactHT100M S3D#3mWAP: 30.5
action-segmentation-on-coinMIL-NCE#6Frame accuracy: 61.0
action-segmentation-on-coinCBT#8Frame accuracy: 53.9
long-video-retrieval-background-removed-onMIL-NCE#6Cap. Avg. R@1: 43.1Cap. Avg. R@5: 68.6Cap. Avg. R@10: 79.1
zero-shot-video-retrieval-on-msr-vttMIL-NCE#32text-to-video R@1: 9.9text-to-video R@5: 24.0text-to-video R@10: 32.4
zero-shot-video-retrieval-on-youcook2MIL-NCE#4text-to-video R@1: 15.1text-to-video R@5: 38.0