Photorealistic Video Generation with Diffusion Models

Benchmark Model Rank Results
text-to-video-generation-on-ucf-101W.A.L.T 3BFVD16: 258.1
video-generation-on-kinetics-600-12-framesW.A.L.T-LFVD: 3.3±0.0
video-generation-on-ucf-101W.A.L.T 3B (text-conditional)FVD16: 258.1Inception Score: 35.1
video-generation-on-ucf-101W.A.L.T-XL (class-conditional)FVD16: 36±2
video-prediction-on-kinetics-600-12-framesW.A.L.T.-LFVD: 3.3