Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation
Open paper
Benchmark
Model
Rank
Results
text-to-video-generation-on-msr-vtt
HiGen
#2
FVD: 406
CLIPSIM: 0.2947
FID: 8.60
Rank counts only results with a code link.