TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Benchmark Model Rank Results
video-question-answering-on-activitynet-qaTESTA (ViT-B/16)#13Accuracy: 45
video-retrieval-on-activitynetTESTA (ViT-B/16)#11text-to-video R@1: 54.8text-to-video R@5: 80.8
video-retrieval-on-condensed-moviesTESTA (ViT-B/16)#1text-to-video R@1: 24.9text-to-video R@5: 46.5
video-retrieval-on-didemoTESTA (ViT-B/16)#9text-to-video R@1: 61.2text-to-video R@5: 87.2
video-retrieval-on-querydTESTA (ViT-B/16)#1text-to-video R@1: 83.4text-to-video R@5: 93.8