OpenCodePapers

speech-recognition-on-librispeech-test-other

Speech Recognition
Dataset Link
Results over time
Click legend items to toggle metrics. Hover points for model names.
Leaderboard
PaperCodeWord Error Rate (WER)ModelNameReleaseDate
Introducing Cohere-transcribe: state-of-the-art speech recognition2.04CohereLabs/cohere-transcribe-03-20262026-03-26
Boson AI Launches Higgs STT 3 Speech-to-Text Model✓ Link2.05bosonai/higgs-audio-v3-8b-stt-v22026-04-27
Introducing Zoom AI Services2.08zoom/scribe_v12026-03-10
Introducing Resonant-1 and Resonant-1-flash2.08reson8/resonant-12026-04-08
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link2.1nvidia/parakeet-rnnt-1.1b2023-12-27
Introducing Resonant-1 and Resonant-1-flash2.13reson8/resonant-1-flash2026-04-08
✓ Link2.17ibm-granite/granite-speech-4.1-2b2026-04-29
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link2.22nvidia/parakeet-tdt-1.1b2024-01-25
Introducing Universal-3 Pro: A new class of speech language model optimized for Voice AI2.27assemblyai/universal-3-pro2026-02-03
Introducing Scribe v22.32elevenlabs/scribe_v22026-01-09
Data-Efficient On-Policy Distillation for Automatic Speech Recognition✓ Link2.35AutoArk-AI/ARK-ASR-3B2026-06-22
NLE: Non-autoregressive LLM-based ASR by Transcript Editing✓ Link2.4ibm-granite/granite-speech-4.1-2b-nar2026-04-01
Kimi-Audio Technical Report✓ Link2.42Kimi-Audio2025-04-25
Step-Audio 2 Technical Report2.42Step-Audio 22025-08-27
Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models2.48SAMBA ASR2025-01-06
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information✓ Link2.49FAdam2024-05-21
✓ Link2.49ibm-granite/granite-4.0-1b-speech2026-03-06
Azure Speech at Build 2026: Powering Voice Agents with Real-Time and Life-like Experiences2.49microsoft/azure-speech-06-20262026-06-04
W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training✓ Link2.5w2v-BERT XXL2021-08-07
Training and Inference Efficiency of Encoder-Decoder Speech Models✓ Link2.52nvidia/canary-1b-flash2025-03-07
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities✓ Link2.52ibm-granite/granite-speech-3.3-8b2025-06-19
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data✓ Link2.53nvidia/canary-1b2024-02-07
Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition✓ Link2.6Conformer + Wav2vec 2.0 + SpecAugment-based Noisy Student Training with Libri-Light2020-10-20
Boson AI Launches Higgs STT 3 Speech-to-Text Model✓ Link2.6bosonai/higgs-audio-v3-stt2026-03-18
✓ Link2.62nvidia/canary-qwen-2.5b2025-07-17
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link2.64nvidia/parakeet-rnnt-0.6b2023-12-28
Pulse STT: Fast, Accurate Speech-to-Text2.7smallestai/pulse2026-01-28
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link2.73nvidia/parakeet-tdt-0.6b-v22025-05-01
Voxtral✓ Link2.79mistralai/Voxtral-Small-24B-25072025-07-15
2.8OpenMOSS-Team/MOSS-Transcribe-preview-2B2026-06-26
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities✓ Link2.84ibm-granite/granite-speech-3.3-2b2025-06-19
Step-Audio 2 Technical Report✓ Link2.86Step-Audio 2 mini2025-08-27
Avalon: ASR for Human–AI Interaction2.87aquavoice/avalon-v1-en2025-08-22
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units✓ Link2.9HuBERT with Libri-Light2021-06-14
Qwen3-ASR Technical Report✓ Link2.92Qwen/Qwen3-ASR-1.7B2026-01-28
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link3.0wav2vec 2.0 with Libri-Light2020-06-20
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units✓ Link3.02facebook/hubert-xlarge-ls960-ft2022-03-02
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST✓ Link3.07nvidia/canary-1b-v22025-08-14
Self-training and Pre-training are Complementary for Speech Recognition✓ Link3.1Conv + Transformer + wav2vec2.0 + pseudo labeling2020-10-22
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST✓ Link3.12nvidia/parakeet-tdt-0.6b-v32025-08-14
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing✓ Link3.2WavLM Large2021-10-26
fairseq S2T: Fast Speech-to-Text Modeling with fairseq✓ Link3.21facebook/wav2vec2-conformer-rel-pos-large-960h-ft2022-04-18
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link3.23nvidia/parakeet-ctc-1.1b2023-12-28
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages✓ Link3.25facebook/omniASR-LLM-7B-v22025-12-12
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network3.3SpeechStew (1B)2021-04-05
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link3.33facebook/wav2vec2-large-960h-lv60-self2020-06-20
SpeechBrain: A General-Purpose Speech Toolkit✓ Link3.36speechbrain/asr-wav2vec2-librispeech2022-06-05
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link3.39nvidia/stt_en_fastconformer_transducer_large2023-06-08
Improved Noisy Student Training for Automatic Speech Recognition✓ Link3.4ContextNet + SpecAugment-based Noisy Student Training with Libri-Light2020-05-19
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link3.41nvidia/parakeet-ctc-0.6b2023-12-28
Training and Inference Efficiency of Encoder-Decoder Speech Models✓ Link3.41nvidia/canary-180m-flash2025-03-11
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language✓ Link3.44facebook/data2vec-audio-large-960h2022-04-02
Data-Efficient On-Policy Distillation for Automatic Speech Recognition✓ Link3.44AutoArk-AI/ARK-ASR-0.6B2026-05-25
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs3.45microsoft/Phi-4-multimodal-instruct2025-02-24
fairseq S2T: Fast Speech-to-Text Modeling with fairseq✓ Link3.46facebook/wav2vec2-conformer-rope-large-960h-ft2022-04-18
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link3.5efficient-speech/lite-whisper-large-v3-acc2025-02-26
Robust Speech Recognition via Large-Scale Weak Supervision✓ Link3.52openai/whisper-large-v32023-11-06
Voxtral✓ Link3.62mistralai/Voxtral-Mini-3B-25072025-07-01
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link3.64nvidia/stt_en_fastconformer_ctc_large2023-06-08
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units✓ Link3.65facebook/hubert-large-ls960-ft2022-03-02
E-Branchformer: Branchformer with Enhanced merging for speech recognition✓ Link3.65E-Branchformer (L) + Internal Language Model Estimation2022-09-30
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language✓ Link3.7data2vec2022-02-07
Robust Speech Recognition via Large-Scale Weak Supervision✓ Link3.7openai/whisper-large-v3-turbo2024-10-01
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link3.73nvidia/stt_en_conformer_ctc_large2022-04-09
GLM-ASR-Nano: A robust, open-source speech recognition model✓ Link3.82zai-org/GLM-ASR-Nano-25122025-12-09
Iterative Pseudo-Labeling for Speech Recognition✓ Link3.83Conv + Transformer AM + Iterative Pseudo-Labeling (n-gram LM + Transformer Rescoring)2020-05-19
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link3.87efficient-speech/lite-whisper-large-v32025-02-26
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link3.9Conformer(L)2020-05-16
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions✓ Link3.95nyrahealth/CrisperWhisper2024-08-29
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link3.95Zipformer+pruned transducer w/ CR-CTC (no external language model)2024-10-07
Qwen3-ASR Technical Report✓ Link3.97Qwen/Qwen3-ASR-0.6B2026-01-28
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling✓ Link3.99kyutai/stt-2.6b-en2025-06-06
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network4.0SpeechStew (100M)2021-04-05
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link4.08efficient-speech/lite-whisper-large-v3-turbo-acc2025-02-26
ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context✓ Link4.1ContextNet(L)2020-05-07
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link4.1wav2vec 2.02020-06-20
End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures✓ Link4.11Conv + Transformer AM (ConvLM with Transformer Rescoring)2019-11-19
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use✓ Link4.12speechbrain/asr-conformer-largescaleasr2025-02-06
Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces4.20CTC + Transformer LM rescoring2020-05-19
Improving RNN Transducer Based ASR with Auxiliary Tasks✓ Link4.20Transformer Transducer2020-11-05
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models✓ Link4.2Qwen-Audio2023-11-14
Step-Audio 2 Technical Report4.23GPT-4o Transcribe2025-08-27
Ursa 2: Elevating speech recognition across 50+ languages4.26speechmatics/enhanced2024-10-11
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link4.3Conformer(M)2020-05-16
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition✓ Link4.34nvidia/nemotron-speech-streaming-en-0.6b2026-03-13
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link4.35Zipformer+CR-CTC (no external language model)2024-10-07
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning✓ Link4.37espnet/owsm_ctc_v4_1B2025-01-16
Zipformer: A faster and better encoder for automatic speech recognition✓ Link4.38Zipformer+pruned transducer (no external language model)2023-10-17
ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition4.46Multistream CNN with Self-Attentive SRU2020-05-21
ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context✓ Link4.5ContextNet(M)2020-05-07
Robust Knowledge Distillation via Large-Scale Pseudo Labelling✓ Link4.5distil-whisper/distil-large-v3.52024-12-05
Zipformer: A faster and better encoder for automatic speech recognition✓ Link4.54soundsgoodai/Zipformer-transducer-XL-290M2026-05-12
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link4.6usefulsensors/moonshine-streaming-medium2026-02-12
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification✓ Link4.65espnet/owsm_ctc_v3.1_1B2024-02-23
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link4.67efficient-speech/lite-whisper-large-v3-fast2025-02-26
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link4.71nvidia/parakeet-tdt_ctc-110m2024-09-17
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training✓ Link4.74facebook/wav2vec2-large-robust-ft-libri-960h2021-04-02
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models✓ Link4.82espnet/owsm_ctc_v3.2_ft_1B2024-09-24
Transformer-based Acoustic Modeling for Hybrid Speech Recognition4.85hybrid + Transformer LM rescoring2019-10-22
Voxtral Realtime✓ Link4.92mistralai/Voxtral-Mini-4B-Realtime-26022026-02-11
VibeVoice-ASR Technical Report✓ Link4.93microsoft/VibeVoice-ASR-HF2026-03-02
Graph Convolutions Enrich the Self-Attention in Transformers!✓ Link4.94Branchformer + GFSA2023-12-07
RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation✓ Link5.0Hybrid model with Transformer rescoring2019-05-08
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link5.0Conformer(S)2020-05-16
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages✓ Link5.05facebook/omniASR-CTC-7B-v22025-12-12
Step-Audio 2 Technical Report5.07Qwen Omni2025-08-27
End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures✓ Link5.18Conv + Transformer AM (ConvLM with Transformer Rescoring) (LS only)2019-11-19
Step-Audio 2 Technical Report5.32Doubao LLM ASR2025-08-27
ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context✓ Link5.5ContextNet(S)2020-05-07
Librispeech Transducer Model with Internal Language Model Prior Correction✓ Link5.6LSTM Transducer2021-04-07
A Comparative Study on Transformer vs RNN in Speech Applications✓ Link5.7Transformer2019-09-13
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition✓ Link5.8LAS + SpecAugment2019-04-18
State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention With Dilated 1D Convolutions✓ Link5.80Multi-Stream Self-Attention With Dilated 1D Convolutions2019-10-01
Squeezeformer: An Efficient Transformer for Automatic Speech Recognition✓ Link5.97Squeezeformer (L)2022-06-02
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link6.02nvidia/stt_en_conformer_transducer_small2022-06-01
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link6.37usefulsensors/moonshine-streaming-small2026-02-12
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language✓ Link6.43facebook/data2vec-audio-base-960h2022-03-02
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition✓ Link6.5LAS (no LM)2019-04-18
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link6.53facebook/wav2vec2-large-960h2020-06-20
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition✓ Link6.78nvidia/nemotron-3.5-asr-streaming-0.6b2026-06-04
Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition✓ Link6.85Conformer with Relaxed Attention2021-07-02
Scaling Speech Technology to 1,000+ Languages✓ Link7.2facebook/mms-1b-all2023-05-27
QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions✓ Link7.25QuartzNet15x52019-10-22
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link7.41nvidia/stt_en_conformer_ctc_small2023-06-12
Moonshine: Speech Recognition for Live Transcription and Voice Commands✓ Link7.62usefulsensors/moonshine-base2024-10-21
Neural Network Language Modeling with Letter-based Features and Importance Sampling7.63tdnn + chain + rnnlm rescoring2018-04-15
Jasper: An End-to-End Convolutional Neural Acoustic Model✓ Link7.84Jasper DR 10x5 (+ Time/Freq Masks)2019-04-05
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link7.95facebook/wav2vec2-base-960h2022-03-02
Espresso: A Fast End-to-end Neural Speech Recognition Toolkit✓ Link8.7Espresso2019-09-18
Jasper: An End-to-End Convolutional Neural Acoustic Model✓ Link8.79Jasper DR 10x52019-04-05
Niagara-38m Sets a New Benchmark for Edge Speech Recognition9.22abr-ai/niagara-38m-batch.en2026-04-15
MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets✓ Link9.6MT4SSL2022-11-14
Fully Convolutional Speech Recognition10.47Convolutional Speech Recognition2018-12-17
CRF-based Single-stage Acoustic Modeling with CTC Topology✓ Link10.65CTC-CRF 4gram-LM2019-04-16
Niagara-38m Sets a New Benchmark for Edge Speech Recognition10.95abr-ai/niagara-19m-batch.en2026-04-15
Moonshine: Speech Recognition for Live Transcription and Voice Commands✓ Link11.06usefulsensors/moonshine-tiny2024-10-21
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link11.54usefulsensors/moonshine-streaming-tiny2026-02-12
12.5TDNN + pNorm + speed up/down speech
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin✓ Link13.25Deep Speech 22015-12-08
Semi-Supervised Speech Recognition via Local Prior Matching✓ Link15.28Local Prior Matching (Large Model, ConvLM LM)2020-02-24
Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces✓ Link16.5Snips2018-05-25
Semi-Supervised Speech Recognition via Local Prior Matching✓ Link20.84Local Prior Matching (Large Model)2020-02-24