How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild

Benchmark Model Rank Results
audio-visual-active-speaker-detection-on-ava-activespeakerASDNet#8validation mean average precision: 93.5%