Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Open paper
Benchmark
Model
Rank
Results
text-to-video-generation-on-evalcrafter-text
Show-1
#2
Visual Quality: 53.74
Motion Quality: 52.19
Temporal Consistency: 60.83
…
text-to-video-generation-on-msr-vtt
Show-1
#4
FVD: 538
CLIPSIM: 0.3072
FID: 13.08
Rank counts only results with a code link.