Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

Benchmark Model Rank Results
highlight-detection-on-tvsumUVCOM (train from scratch)#5mAP: 86.3
highlight-detection-on-youtube-highlightsUVCOM#2mAP: 77.4
moment-retrieval-on-charades-staUVCOM#14R@1 IoU=0.5: 59.25R@1 IoU=0.7: 36.64
moment-retrieval-on-qvhighlightsUVCOM (w/ PT ASR Captions)#16mAP: 43.8R@1 IoU=0.5: 64.53R@1 IoU=0.7: 48.31mAP@0.5: 64.78
moment-retrieval-on-qvhighlightsUVCOM#18mAP: 43.18R@1 IoU=0.5: 63.55R@1 IoU=0.7: 47.47mAP@0.5: 63.37
natural-language-moment-retrieval-on-tacosUVCOM#13R@1,IoU=0.5: 36.39R@1,IoU=0.7: 23.32