COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval
Open paper
Benchmark
Model
Rank
Results
video-retrieval-on-msr-vtt
COTS
–
text-to-video R@1: 32.1
text-to-video R@5: 60.8
…
video-retrieval-on-msr-vtt-1ka
COTS
–
text-to-video R@1: 36.8
text-to-video R@5: 63.8
…
Rank counts only results with a code link.