Tell What You Hear From What You See -- Video to Audio Generation Through Text

Benchmark Model Rank Results
video-to-sound-generation-on-vgg-soundVATT-LLama#7FAD: 2.38KLD: 1.41