Tell What You Hear From What You See -- Video to Audio Generation Through Text
Open paper
Benchmark
Model
Rank
Results
video-to-sound-generation-on-vgg-sound
VATT-LLama
#7
FAD: 2.38
KLD: 1.41
Rank counts only results with a code link.