Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Benchmark Model Rank Results
image-generation-on-imagenet-256x256MAGVIT-v2#36FID: 1.78
image-generation-on-imagenet-256x256MAGVIT-v2 (w/o guidance)#67FID: 3.65
image-generation-on-imagenet-512x512MAGVIT-v2#23FID: 1.91Inception score: 324.3
image-generation-on-imagenet-512x512MAGVIT-v2 (w/o guidance)#39FID: 3.07Inception score: 213.1
video-generation-on-kinetics-600-12-framesMAGVIT-v2#1FVD: 4.3±0.1
video-generation-on-ucf-101MAGVIT-v2#4FVD16: 58±3
video-generation-on-ucf-101MAGVIT-v2 (AR)#8FVD16: 109
video-prediction-on-kinetics-600-12-framesMAGVIT-v2#1FVD: 4.3±0.1