UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Benchmark Model Rank Results
highlight-detection-on-qvhighlightsUMT (w. PT)#12mAP: 39.12
highlight-detection-on-qvhighlightsUMT#17mAP: 38.18
highlight-detection-on-tvsumUMT#7mAP: 83.1
highlight-detection-on-youtube-highlightsUMT#7mAP: 74.9
moment-retrieval-on-charades-staUMT (VO)#22R@1 IoU=0.5: 49.35R@1 IoU=0.7: 26.16R@5 IoU=0.5: 89.41
moment-retrieval-on-charades-staUMT (VA)#24R@1 IoU=0.5: 48.31R@1 IoU=0.7: 29.25R@5 IoU=0.5: 88.79
moment-retrieval-on-qvhighlightsUMT (w/ audio + PT ASR Cpations)#25mAP: 38.08
moment-retrieval-on-qvhighlightsUMT#27mAP: 36.12
video-grounding-on-qvhighlightsUMT#5R@1,IoU=0.7: 41.18R@1,IoU=0.5: 56.23