Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation

Benchmark Model Rank Results
text-to-video-generation-on-msr-vttHiGen#2FVD: 406CLIPSIM: 0.2947FID: 8.60