LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Benchmark Model Rank Results
text-to-video-generation-on-evalcrafter-textLavie#4Visual Quality: 52.83Motion Quality: 57.99Temporal Consistency: 54.23
text-to-video-generation-on-ucf-101LAVIE (Zero-shot, 320x512)#2FVD16: 526.30
video-generation-on-ucf-101LAVIE (320x512, text-conditional)#29FVD16: 526.30