UniVTG: Towards Unified Video-Language Temporal Grounding

Benchmark Model Rank Results
highlight-detection-on-qvhighlightsUniVTG (w/ PT)#8mAP: 40.54Hit@1: 66.28
highlight-detection-on-qvhighlightsUniVTG#16mAP: 38.20Hit@1: 60.96
moment-retrieval-on-qvhighlightsUniVTG (w/ PT)#17mAP: 43.63R@1 IoU=0.5: 65.43R@1 IoU=0.7: 50.06mAP@0.5: 64.06
moment-retrieval-on-qvhighlightsUniVTG#28mAP: 35.47R@1 IoU=0.5: 58.86R@1 IoU=0.7: 40.86mAP@0.5: 57.60
natural-language-moment-retrieval-on-tacosUniVTG#9R@1,IoU=0.3: 51.44R@1,IoU=0.5: 34.97R@1,IoU=0.7: 21.07mIoU: 35.76