iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

Benchmark Model Rank Results
dense-video-captioning-on-activitynet-captionsiPerceive (Chadha et al., 2020)–METEOR: 7.87BLEU-3: 2.93BLEU-4: 1.29
video-question-answering-on-tvqaiPerceive (Chadha et al., 2020)–Accuracy: 76.96