VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Benchmark Model Rank Results
action-segmentation-on-coinVLM#5Frame accuracy: 68.4
temporal-action-localization-on-crosstaskVLM#2Recall: 46.5
video-captioning-on-youcook2VLM#4BLEU-4: 12.27BLEU-3: 17.78CIDEr: 1.3869ROUGE-L: 41.51
video-retrieval-on-msr-vtt-1kaVLM#43text-to-video R@1: 28.10text-to-video R@5: 55.50
video-retrieval-on-youcook2VLM#5text-to-video R@1: 27.05text-to-video R@5: 56.88