VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Benchmark Model Rank Results
video-captioning-on-msr-vtt-1VLABCIDEr: 74.9METEOR: 33.4ROUGE-L: 68.3BLEU-4: 54.6
video-captioning-on-msvd-1VLABCIDEr: 179.8BLEU-4: 79.3METEOR: 51.2ROUGE-L: 87.9
video-retrieval-on-didemoVLABtext-to-video R@1: 56.8text-to-video R@5: 81.6
video-retrieval-on-msr-vttVLABtext-to-video R@1: 55.1text-to-video R@5: 78.8
video-retrieval-on-msvdVLABtext-to-video R@1: 57.5text-to-video R@5: 83.6
visual-question-answering-on-msrvtt-qa-1VLABAccuracy: 0.496
visual-question-answering-on-msvd-qa-1VLABAccuracy: 0.61