Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

Benchmark Model Rank Results
video-question-answering-on-msrvtt-qaHBI#7Accuracy: 46.2
video-retrieval-on-activitynetHBI#21text-to-video R@1: 42.2text-to-video R@5: 73.0…
video-retrieval-on-didemoHBI#24text-to-video R@1: 46.9text-to-video R@5: 74.9…
video-retrieval-on-msr-vtt-1kaHBI#20text-to-video R@1: 48.6text-to-video R@5: 74.6…
visual-question-answering-on-msrvtt-qaHBI#8Accuracy: 0.462