Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems
Open paper
Benchmark
Model
Rank
Results
audio-visual-speech-recognition-on-lrw
AVCRFormer
#1
Top-1 Accuracy: 98.81
lipreading-on-lip-reading-in-the-wild
AVCRFormer
#4
Top-1 Accuracy: 89.57
Rank counts only results with a code link.