VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Benchmark Model Rank Results
audio-captioning-on-audiocapsVALOR#12CIDEr: 0.741BLEU-4: 0.270METEOR: 0.231ROUGE-L: 0.494
audio-captioning-on-clothoVALOR#7CIDEr: 0.423BLEU-4: 16.2METEOR: 17.4ROUGE-L: 38.2
audio-visual-question-answering-on-music-avqaVALOR#2Acc: 78.9
cross-modal-retrieval-on-coco-2014VALOR#12Text-to-image R@1: 61.4Text-to-image R@5: 84.4
image-captioning-on-coco-captionsVALOR#32CIDER: 152.5SPICE: 25.7
text-to-audio-retrieval-on-audiocapsVALOR#4R@1: 40.1R@5: 73.9R@10: 83.1
text-to-audio-retrieval-on-clothoVALOR#6R@1: 17.5R@5: 42.7R@10: 55.3
video-captioning-on-msr-vtt-1VALOR#5CIDEr: 74.0METEOR: 32.9ROUGE-L: 68.0BLEU-4: 54.4
video-captioning-on-msvd-1VALOR#2CIDEr: 178.5BLEU-4: 80.7METEOR: 51.0ROUGE-L: 87.9
video-captioning-on-vatex-1VALOR#1BLEU-4: 45.6CIDEr: 95.8METEOR: 29.4ROUGE-L: 57.4
video-question-answering-on-activitynet-qaVALOR#5Accuracy: 48.6
video-question-answering-on-msrvtt-qaVALOR#2Accuracy: 49.2
video-retrieval-on-activitynetVALOR#3text-to-video R@1: 70.1text-to-video R@5: 90.8
video-retrieval-on-didemoVALOR#7text-to-video R@1: 61.5text-to-video R@5: 85.3
video-retrieval-on-lsmdcVALOR#6text-to-video R@1: 34.2text-to-video R@5: 56.0
video-retrieval-on-msr-vttVALOR#4text-to-video R@1: 59.9text-to-video R@5: 83.5
video-retrieval-on-vatexVALOR#3text-to-video R@1: 78.5text-to-video R@5: 97.1
visual-question-answering-on-msvd-qa-1VALOR#3Accuracy: 0.60
visual-question-answering-on-vqa-v2-test-devVALOR#12Accuracy: 78.46
visual-question-answering-on-vqa-v2-test-stdVALOR#8overall: 78.62