Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
Open paper
Benchmark
Model
Rank
Results
lipreading-on-lip-reading-in-the-wild
3D Conv + ResNet-18 + Bi-GRU + Visual-Audio Memory
#9
Top-1 Accuracy: 85.4
lipreading-on-lrw-1000
3D Conv + ResNet-18 + Bi-GRU + Visual-Audio Memory
#4
Top-1 Accuracy: 50.82%
Rank counts only results with a code link.