Query-Dependent Video Representation for Moment Retrieval and Highlight Detection

Benchmark Model Rank Results
highlight-detection-on-qvhighlightsQD-DETR#13mAP: 39.04Hit@1: 62.87
highlight-detection-on-qvhighlightsQD-DETR (only Video)#14mAP: 38.94Hit@1: 62.40
highlight-detection-on-qvhighlightsQD-DETR (w/ PT)#15mAP: 38.52Hit@1: 62.27
highlight-detection-on-qvhighlightsQD-DETR (only Video w/ PT)#20Hit@1: 61.91
highlight-detection-on-tvsumQD-DETR#4mAP: 86.6
highlight-detection-on-tvsumQD-DETR (only Video)#6mAP: 85.0
moment-retrieval-on-charades-staQD-DETR (Only Video)#18R@1 IoU=0.5: 57.31R@1 IoU=0.7: 32.55
moment-retrieval-on-qvhighlightsQD-DETR (w/ PT)#20mAP: 40.62R@1 IoU=0.5: 64.1R@1 IoU=0.7: 46.1mAP@0.5: 64.3
moment-retrieval-on-qvhighlightsQD-DETR (w/ audio)#21mAP: 40.19R@1 IoU=0.5: 63.06R@1 IoU=0.7: 45.10mAP@0.5: 63.04
moment-retrieval-on-qvhighlightsQD-DETR (only Video w/ PT ASR Captions)#23mAP: 40.0R@1 IoU=0.5: 63.2R@1 IoU=0.7: 45.2mAP@0.5: 63.4
moment-retrieval-on-qvhighlightsQD-DETR (only Video)#24mAP: 39.86R@1 IoU=0.5: 62.40R@1 IoU=0.7: 44.98mAP@0.5: 62.52
video-grounding-on-qvhighlightsQD-DETR#4R@1,IoU=0.7: 44.98R@1,IoU=0.5: 62.40