How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild
Open paper
Benchmark
Model
Rank
Results
audio-visual-active-speaker-detection-on-ava-activespeaker
ASDNet
#8
validation mean average precision: 93.5%
Rank counts only results with a code link.