CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter

Benchmark Model Rank Results
video-captioning-on-msr-vtt-1CLIP-DCD#10CIDEr: 58.7METEOR: 31.3ROUGE-L: 64.8BLEU-4: 48.2