CenterCLIP: Token Clustering for Efficient Text-Video Retrieval

Benchmark Model Rank Results
video-retrieval-on-activitynetCenterCLIP (ViT-B/16)#18text-to-video R@1: 46.2text-to-video R@5: 77.0
video-retrieval-on-lsmdcCenterCLIP (ViT-B/16)#16text-to-video R@1: 24.2text-to-video R@5: 46.2
video-retrieval-on-msr-vtt-1kaCenterCLIP (ViT-B/16)#21text-to-video R@1: 48.4text-to-video R@5: 73.8
video-retrieval-on-msvdCenterCLIP (ViT-B/16)#7text-to-video R@1: 50.6text-to-video R@5: 80.3