A Recipe for Scaling up Text-to-Video Generation with Text-free Videos
Open paper
Benchmark
Model
Rank
Results
text-to-video-generation-on-msr-vtt
TF-T2V
#3
FVD: 441
CLIPSIM: 0.2991
FID: 8.19
Rank counts only results with a code link.