| audio-captioning-on-audiocaps | VALOR | #12 | CIDEr: 0.741BLEU-4: 0.270METEOR: 0.231ROUGE-L: 0.494 |
| audio-captioning-on-clotho | VALOR | #7 | CIDEr: 0.423BLEU-4: 16.2METEOR: 17.4ROUGE-L: 38.2 |
| audio-visual-question-answering-on-music-avqa | VALOR | #2 | Acc: 78.9 |
| cross-modal-retrieval-on-coco-2014 | VALOR | #12 | Text-to-image R@1: 61.4Text-to-image R@5: 84.4… |
| image-captioning-on-coco-captions | VALOR | #32 | CIDER: 152.5SPICE: 25.7 |
| text-to-audio-retrieval-on-audiocaps | VALOR | #4 | R@1: 40.1R@5: 73.9R@10: 83.1 |
| text-to-audio-retrieval-on-clotho | VALOR | #6 | R@1: 17.5R@5: 42.7R@10: 55.3 |
| video-captioning-on-msr-vtt-1 | VALOR | #5 | CIDEr: 74.0METEOR: 32.9ROUGE-L: 68.0BLEU-4: 54.4 |
| video-captioning-on-msvd-1 | VALOR | #2 | CIDEr: 178.5BLEU-4: 80.7METEOR: 51.0ROUGE-L: 87.9 |
| video-captioning-on-vatex-1 | VALOR | #1 | BLEU-4: 45.6CIDEr: 95.8METEOR: 29.4ROUGE-L: 57.4 |
| video-question-answering-on-activitynet-qa | VALOR | #5 | Accuracy: 48.6 |
| video-question-answering-on-msrvtt-qa | VALOR | #2 | Accuracy: 49.2 |
| video-retrieval-on-activitynet | VALOR | #3 | text-to-video R@1: 70.1text-to-video R@5: 90.8… |
| video-retrieval-on-didemo | VALOR | #7 | text-to-video R@1: 61.5text-to-video R@5: 85.3… |
| video-retrieval-on-lsmdc | VALOR | #6 | text-to-video R@1: 34.2text-to-video R@5: 56.0… |
| video-retrieval-on-msr-vtt | VALOR | #4 | text-to-video R@1: 59.9text-to-video R@5: 83.5… |
| video-retrieval-on-vatex | VALOR | #3 | text-to-video R@1: 78.5text-to-video R@5: 97.1… |
| visual-question-answering-on-msvd-qa-1 | VALOR | #3 | Accuracy: 0.60 |
| visual-question-answering-on-vqa-v2-test-dev | VALOR | #12 | Accuracy: 78.46 |
| visual-question-answering-on-vqa-v2-test-std | VALOR | #8 | overall: 78.62 |