Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Benchmark Model Rank Results
text-to-video-generation-on-evalcrafter-textShow-1#2Visual Quality: 53.74Motion Quality: 52.19Temporal Consistency: 60.83
text-to-video-generation-on-msr-vttShow-1#4FVD: 538CLIPSIM: 0.3072FID: 13.08