Audio Retrieval with Natural Language Queries: A Benchmark Study

Benchmark Model Rank Results
text-to-audio-retrieval-on-audiocapsMMT#5R@1: 36.1±3.3R@10: 84.5±2.0
text-to-audio-retrieval-on-audiocapsCE#7R@1: 23.6± 0.6R@10: 71.4±0.5
text-to-audio-retrieval-on-audiocapsMoEE#9R@1: 23.0±0.7R@10: 71.0±1.2
text-to-audio-retrieval-on-clothoMMT#10R@1: 6.5±0.6R@10: 32.8±2.1
text-to-audio-retrieval-on-clothoCE(pretraining:SoundDescs)#11R@1: 6.4±0.5R@10: 32.5±1.7
text-to-audio-retrieval-on-sounddescsCE#1R@1: 31.1±0.2R@10: 70.8±0.5
text-to-audio-retrieval-on-sounddescsMoEE#2R@1: 30.8±0.7R@10: 70.9±0.5
text-to-audio-retrieval-on-sounddescsMMT#3R@1: 30.7±0.4R@10: 72.7±0.8
text-to-audio-retrieval-on-sounddescsCE(pretrained: AudioCaps)#4R@1: 23.3±0.7R@10: 63.9±0.5