End-to-end Generative Pretraining for Multimodal Video Captioning
Open paper
Benchmark
Model
Rank
Results
video-captioning-on-msr-vtt-1
MV-GPT
–
CIDEr: 60.0
METEOR: 38.7
ROUGE-L: 64.0
BLEU-4: 48.9
Rank counts only results with a code link.