Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

Benchmark Model Rank Results
science-question-answering-on-scienceqaVideo-LaVIT#9Avg. Accuracy: 70.0
text-to-video-generation-on-msr-vttVideo-LaVIT#1FVD: 188.36CLIPSIM: 0.3012FID: 11.27
video-generation-on-ucf-101Video-LaVIT#16FVD16: 280.57Inception Score: 44.26
visual-question-answering-on-gqa-test-devVideo-LaVIT#3Accuracy: 64.4
visual-question-answering-on-mm-vetVideo-LaVIT#122GPT-4 score: 33.2Params: 7B
visual-question-answering-on-mmbenchVideo-LaVIT#4GPT-3.5 score: 67.3
visual-question-answering-on-vizwiz-2020-vqaVideo-LaVIT#2overall: 56.0
zeroshot-video-question-answer-on-activitynetVideo-LaVIT#13Accuracy: 50.1Confidence Score: 3.3
zeroshot-video-question-answer-on-msrvtt-qaVideo-LaVIT#15Accuracy: 59.3Confidence Score: 3.3
zeroshot-video-question-answer-on-msvd-qaVideo-LaVIT#14Accuracy: 73.2Confidence Score: 3.9