| Introducing Cohere-transcribe: state-of-the-art speech recognition | | 2.04 | CohereLabs/cohere-transcribe-03-2026 | 2026-03-26 |
| Boson AI Launches Higgs STT 3 Speech-to-Text Model | ✓ Link | 2.05 | bosonai/higgs-audio-v3-8b-stt-v2 | 2026-04-27 |
| Introducing Zoom AI Services | | 2.08 | zoom/scribe_v1 | 2026-03-10 |
| Introducing Resonant-1 and Resonant-1-flash | | 2.08 | reson8/resonant-1 | 2026-04-08 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 2.1 | nvidia/parakeet-rnnt-1.1b | 2023-12-27 |
| Introducing Resonant-1 and Resonant-1-flash | | 2.13 | reson8/resonant-1-flash | 2026-04-08 |
| ✓ Link | 2.17 | ibm-granite/granite-speech-4.1-2b | 2026-04-29 |
| Efficient Sequence Transduction by Jointly Predicting Tokens and Durations | ✓ Link | 2.22 | nvidia/parakeet-tdt-1.1b | 2024-01-25 |
| Introducing Universal-3 Pro: A new class of speech language model optimized for Voice AI | | 2.27 | assemblyai/universal-3-pro | 2026-02-03 |
| Introducing Scribe v2 | | 2.32 | elevenlabs/scribe_v2 | 2026-01-09 |
| Data-Efficient On-Policy Distillation for Automatic Speech Recognition | ✓ Link | 2.35 | AutoArk-AI/ARK-ASR-3B | 2026-06-22 |
| NLE: Non-autoregressive LLM-based ASR by Transcript Editing | ✓ Link | 2.4 | ibm-granite/granite-speech-4.1-2b-nar | 2026-04-01 |
| Kimi-Audio Technical Report | ✓ Link | 2.42 | Kimi-Audio | 2025-04-25 |
| Step-Audio 2 Technical Report | | 2.42 | Step-Audio 2 | 2025-08-27 |
| Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models | | 2.48 | SAMBA ASR | 2025-01-06 |
| FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information | ✓ Link | 2.49 | FAdam | 2024-05-21 |
| ✓ Link | 2.49 | ibm-granite/granite-4.0-1b-speech | 2026-03-06 |
| Azure Speech at Build 2026: Powering Voice Agents with Real-Time and Life-like Experiences | | 2.49 | microsoft/azure-speech-06-2026 | 2026-06-04 |
| W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training | ✓ Link | 2.5 | w2v-BERT XXL | 2021-08-07 |
| Training and Inference Efficiency of Encoder-Decoder Speech Models | ✓ Link | 2.52 | nvidia/canary-1b-flash | 2025-03-07 |
| Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities | ✓ Link | 2.52 | ibm-granite/granite-speech-3.3-8b | 2025-06-19 |
| Less is More: Accurate Speech Recognition & Translation without Web-Scale Data | ✓ Link | 2.53 | nvidia/canary-1b | 2024-02-07 |
| Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition | ✓ Link | 2.6 | Conformer + Wav2vec 2.0 + SpecAugment-based Noisy Student Training with Libri-Light | 2020-10-20 |
| Boson AI Launches Higgs STT 3 Speech-to-Text Model | ✓ Link | 2.6 | bosonai/higgs-audio-v3-stt | 2026-03-18 |
| ✓ Link | 2.62 | nvidia/canary-qwen-2.5b | 2025-07-17 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 2.64 | nvidia/parakeet-rnnt-0.6b | 2023-12-28 |
| Pulse STT: Fast, Accurate Speech-to-Text | | 2.7 | smallestai/pulse | 2026-01-28 |
| Efficient Sequence Transduction by Jointly Predicting Tokens and Durations | ✓ Link | 2.73 | nvidia/parakeet-tdt-0.6b-v2 | 2025-05-01 |
| Voxtral | ✓ Link | 2.79 | mistralai/Voxtral-Small-24B-2507 | 2025-07-15 |
| | 2.8 | OpenMOSS-Team/MOSS-Transcribe-preview-2B | 2026-06-26 |
| Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities | ✓ Link | 2.84 | ibm-granite/granite-speech-3.3-2b | 2025-06-19 |
| Step-Audio 2 Technical Report | ✓ Link | 2.86 | Step-Audio 2 mini | 2025-08-27 |
| Avalon: ASR for Human–AI Interaction | | 2.87 | aquavoice/avalon-v1-en | 2025-08-22 |
| HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units | ✓ Link | 2.9 | HuBERT with Libri-Light | 2021-06-14 |
| Qwen3-ASR Technical Report | ✓ Link | 2.92 | Qwen/Qwen3-ASR-1.7B | 2026-01-28 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | ✓ Link | 3.0 | wav2vec 2.0 with Libri-Light | 2020-06-20 |
| HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units | ✓ Link | 3.02 | facebook/hubert-xlarge-ls960-ft | 2022-03-02 |
| Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST | ✓ Link | 3.07 | nvidia/canary-1b-v2 | 2025-08-14 |
| Self-training and Pre-training are Complementary for Speech Recognition | ✓ Link | 3.1 | Conv + Transformer + wav2vec2.0 + pseudo labeling | 2020-10-22 |
| Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST | ✓ Link | 3.12 | nvidia/parakeet-tdt-0.6b-v3 | 2025-08-14 |
| WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing | ✓ Link | 3.2 | WavLM Large | 2021-10-26 |
| fairseq S2T: Fast Speech-to-Text Modeling with fairseq | ✓ Link | 3.21 | facebook/wav2vec2-conformer-rel-pos-large-960h-ft | 2022-04-18 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 3.23 | nvidia/parakeet-ctc-1.1b | 2023-12-28 |
| Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages | ✓ Link | 3.25 | facebook/omniASR-LLM-7B-v2 | 2025-12-12 |
| SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network | | 3.3 | SpeechStew (1B) | 2021-04-05 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | ✓ Link | 3.33 | facebook/wav2vec2-large-960h-lv60-self | 2020-06-20 |
| SpeechBrain: A General-Purpose Speech Toolkit | ✓ Link | 3.36 | speechbrain/asr-wav2vec2-librispeech | 2022-06-05 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 3.39 | nvidia/stt_en_fastconformer_transducer_large | 2023-06-08 |
| Improved Noisy Student Training for Automatic Speech Recognition | ✓ Link | 3.4 | ContextNet + SpecAugment-based Noisy Student Training with Libri-Light | 2020-05-19 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 3.41 | nvidia/parakeet-ctc-0.6b | 2023-12-28 |
| Training and Inference Efficiency of Encoder-Decoder Speech Models | ✓ Link | 3.41 | nvidia/canary-180m-flash | 2025-03-11 |
| data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language | ✓ Link | 3.44 | facebook/data2vec-audio-large-960h | 2022-04-02 |
| Data-Efficient On-Policy Distillation for Automatic Speech Recognition | ✓ Link | 3.44 | AutoArk-AI/ARK-ASR-0.6B | 2026-05-25 |
| Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs | | 3.45 | microsoft/Phi-4-multimodal-instruct | 2025-02-24 |
| fairseq S2T: Fast Speech-to-Text Modeling with fairseq | ✓ Link | 3.46 | facebook/wav2vec2-conformer-rope-large-960h-ft | 2022-04-18 |
| LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation | ✓ Link | 3.5 | efficient-speech/lite-whisper-large-v3-acc | 2025-02-26 |
| Robust Speech Recognition via Large-Scale Weak Supervision | ✓ Link | 3.52 | openai/whisper-large-v3 | 2023-11-06 |
| Voxtral | ✓ Link | 3.62 | mistralai/Voxtral-Mini-3B-2507 | 2025-07-01 |
| Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition | ✓ Link | 3.64 | nvidia/stt_en_fastconformer_ctc_large | 2023-06-08 |
| HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units | ✓ Link | 3.65 | facebook/hubert-large-ls960-ft | 2022-03-02 |
| E-Branchformer: Branchformer with Enhanced merging for speech recognition | ✓ Link | 3.65 | E-Branchformer (L) + Internal Language Model Estimation | 2022-09-30 |
| data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language | ✓ Link | 3.7 | data2vec | 2022-02-07 |
| Robust Speech Recognition via Large-Scale Weak Supervision | ✓ Link | 3.7 | openai/whisper-large-v3-turbo | 2024-10-01 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 3.73 | nvidia/stt_en_conformer_ctc_large | 2022-04-09 |
| GLM-ASR-Nano: A robust, open-source speech recognition model | ✓ Link | 3.82 | zai-org/GLM-ASR-Nano-2512 | 2025-12-09 |
| Iterative Pseudo-Labeling for Speech Recognition | ✓ Link | 3.83 | Conv + Transformer AM + Iterative Pseudo-Labeling (n-gram LM + Transformer Rescoring) | 2020-05-19 |
| LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation | ✓ Link | 3.87 | efficient-speech/lite-whisper-large-v3 | 2025-02-26 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 3.9 | Conformer(L) | 2020-05-16 |
| CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions | ✓ Link | 3.95 | nyrahealth/CrisperWhisper | 2024-08-29 |
| CR-CTC: Consistency regularization on CTC for improved speech recognition | ✓ Link | 3.95 | Zipformer+pruned transducer w/ CR-CTC
(no external language model) | 2024-10-07 |
| Qwen3-ASR Technical Report | ✓ Link | 3.97 | Qwen/Qwen3-ASR-0.6B | 2026-01-28 |
| Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling | ✓ Link | 3.99 | kyutai/stt-2.6b-en | 2025-06-06 |
| SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network | | 4.0 | SpeechStew (100M) | 2021-04-05 |
| LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation | ✓ Link | 4.08 | efficient-speech/lite-whisper-large-v3-turbo-acc | 2025-02-26 |
| ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context | ✓ Link | 4.1 | ContextNet(L) | 2020-05-07 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | ✓ Link | 4.1 | wav2vec 2.0 | 2020-06-20 |
| End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures | ✓ Link | 4.11 | Conv + Transformer AM (ConvLM with Transformer Rescoring) | 2019-11-19 |
| Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use | ✓ Link | 4.12 | speechbrain/asr-conformer-largescaleasr | 2025-02-06 |
| Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces | | 4.20 | CTC + Transformer LM rescoring | 2020-05-19 |
| Improving RNN Transducer Based ASR with Auxiliary Tasks | ✓ Link | 4.20 | Transformer Transducer | 2020-11-05 |
| Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models | ✓ Link | 4.2 | Qwen-Audio | 2023-11-14 |
| Step-Audio 2 Technical Report | | 4.23 | GPT-4o Transcribe | 2025-08-27 |
| Ursa 2: Elevating speech recognition across 50+ languages | | 4.26 | speechmatics/enhanced | 2024-10-11 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 4.3 | Conformer(M) | 2020-05-16 |
| Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition | ✓ Link | 4.34 | nvidia/nemotron-speech-streaming-en-0.6b | 2026-03-13 |
| CR-CTC: Consistency regularization on CTC for improved speech recognition | ✓ Link | 4.35 | Zipformer+CR-CTC
(no external language model) | 2024-10-07 |
| OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning | ✓ Link | 4.37 | espnet/owsm_ctc_v4_1B | 2025-01-16 |
| Zipformer: A faster and better encoder for automatic speech recognition | ✓ Link | 4.38 | Zipformer+pruned transducer
(no external language model) | 2023-10-17 |
| ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition | | 4.46 | Multistream CNN with Self-Attentive SRU | 2020-05-21 |
| ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context | ✓ Link | 4.5 | ContextNet(M) | 2020-05-07 |
| Robust Knowledge Distillation via Large-Scale Pseudo Labelling | ✓ Link | 4.5 | distil-whisper/distil-large-v3.5 | 2024-12-05 |
| Zipformer: A faster and better encoder for automatic speech recognition | ✓ Link | 4.54 | soundsgoodai/Zipformer-transducer-XL-290M | 2026-05-12 |
| Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications | ✓ Link | 4.6 | usefulsensors/moonshine-streaming-medium | 2026-02-12 |
| OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification | ✓ Link | 4.65 | espnet/owsm_ctc_v3.1_1B | 2024-02-23 |
| LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation | ✓ Link | 4.67 | efficient-speech/lite-whisper-large-v3-fast | 2025-02-26 |
| Efficient Sequence Transduction by Jointly Predicting Tokens and Durations | ✓ Link | 4.71 | nvidia/parakeet-tdt_ctc-110m | 2024-09-17 |
| Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training | ✓ Link | 4.74 | facebook/wav2vec2-large-robust-ft-libri-960h | 2021-04-02 |
| On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models | ✓ Link | 4.82 | espnet/owsm_ctc_v3.2_ft_1B | 2024-09-24 |
| Transformer-based Acoustic Modeling for Hybrid Speech Recognition | | 4.85 | hybrid + Transformer LM rescoring | 2019-10-22 |
| Voxtral Realtime | ✓ Link | 4.92 | mistralai/Voxtral-Mini-4B-Realtime-2602 | 2026-02-11 |
| VibeVoice-ASR Technical Report | ✓ Link | 4.93 | microsoft/VibeVoice-ASR-HF | 2026-03-02 |
| Graph Convolutions Enrich the Self-Attention in Transformers! | ✓ Link | 4.94 | Branchformer + GFSA | 2023-12-07 |
| RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation | ✓ Link | 5.0 | Hybrid model with Transformer rescoring | 2019-05-08 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 5.0 | Conformer(S) | 2020-05-16 |
| Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages | ✓ Link | 5.05 | facebook/omniASR-CTC-7B-v2 | 2025-12-12 |
| Step-Audio 2 Technical Report | | 5.07 | Qwen Omni | 2025-08-27 |
| End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures | ✓ Link | 5.18 | Conv + Transformer AM (ConvLM with Transformer Rescoring) (LS only) | 2019-11-19 |
| Step-Audio 2 Technical Report | | 5.32 | Doubao LLM ASR | 2025-08-27 |
| ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context | ✓ Link | 5.5 | ContextNet(S) | 2020-05-07 |
| Librispeech Transducer Model with Internal Language Model Prior Correction | ✓ Link | 5.6 | LSTM Transducer | 2021-04-07 |
| A Comparative Study on Transformer vs RNN in Speech Applications | ✓ Link | 5.7 | Transformer | 2019-09-13 |
| SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition | ✓ Link | 5.8 | LAS + SpecAugment | 2019-04-18 |
| State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention With Dilated 1D Convolutions | ✓ Link | 5.80 | Multi-Stream Self-Attention With Dilated 1D Convolutions | 2019-10-01 |
| Squeezeformer: An Efficient Transformer for Automatic Speech Recognition | ✓ Link | 5.97 | Squeezeformer (L) | 2022-06-02 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 6.02 | nvidia/stt_en_conformer_transducer_small | 2022-06-01 |
| Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications | ✓ Link | 6.37 | usefulsensors/moonshine-streaming-small | 2026-02-12 |
| data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language | ✓ Link | 6.43 | facebook/data2vec-audio-base-960h | 2022-03-02 |
| SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition | ✓ Link | 6.5 | LAS (no LM) | 2019-04-18 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | ✓ Link | 6.53 | facebook/wav2vec2-large-960h | 2020-06-20 |
| Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition | ✓ Link | 6.78 | nvidia/nemotron-3.5-asr-streaming-0.6b | 2026-06-04 |
| Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition | ✓ Link | 6.85 | Conformer with Relaxed Attention | 2021-07-02 |
| Scaling Speech Technology to 1,000+ Languages | ✓ Link | 7.2 | facebook/mms-1b-all | 2023-05-27 |
| QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions | ✓ Link | 7.25 | QuartzNet15x5 | 2019-10-22 |
| Conformer: Convolution-augmented Transformer for Speech Recognition | ✓ Link | 7.41 | nvidia/stt_en_conformer_ctc_small | 2023-06-12 |
| Moonshine: Speech Recognition for Live Transcription and Voice Commands | ✓ Link | 7.62 | usefulsensors/moonshine-base | 2024-10-21 |
| Neural Network Language Modeling with Letter-based Features and Importance Sampling | | 7.63 | tdnn + chain + rnnlm rescoring | 2018-04-15 |
| Jasper: An End-to-End Convolutional Neural Acoustic Model | ✓ Link | 7.84 | Jasper DR 10x5 (+ Time/Freq Masks) | 2019-04-05 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | ✓ Link | 7.95 | facebook/wav2vec2-base-960h | 2022-03-02 |
| Espresso: A Fast End-to-end Neural Speech Recognition Toolkit | ✓ Link | 8.7 | Espresso | 2019-09-18 |
| Jasper: An End-to-End Convolutional Neural Acoustic Model | ✓ Link | 8.79 | Jasper DR 10x5 | 2019-04-05 |
| Niagara-38m Sets a New Benchmark for Edge Speech Recognition | | 9.22 | abr-ai/niagara-38m-batch.en | 2026-04-15 |
| MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets | ✓ Link | 9.6 | MT4SSL | 2022-11-14 |
| Fully Convolutional Speech Recognition | | 10.47 | Convolutional Speech Recognition | 2018-12-17 |
| CRF-based Single-stage Acoustic Modeling with CTC Topology | ✓ Link | 10.65 | CTC-CRF 4gram-LM | 2019-04-16 |
| Niagara-38m Sets a New Benchmark for Edge Speech Recognition | | 10.95 | abr-ai/niagara-19m-batch.en | 2026-04-15 |
| Moonshine: Speech Recognition for Live Transcription and Voice Commands | ✓ Link | 11.06 | usefulsensors/moonshine-tiny | 2024-10-21 |
| Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications | ✓ Link | 11.54 | usefulsensors/moonshine-streaming-tiny | 2026-02-12 |
| | 12.5 | TDNN + pNorm + speed up/down speech | |
| Deep Speech 2: End-to-End Speech Recognition in English and Mandarin | ✓ Link | 13.25 | Deep Speech 2 | 2015-12-08 |
| Semi-Supervised Speech Recognition via Local Prior Matching | ✓ Link | 15.28 | Local Prior Matching (Large Model, ConvLM LM) | 2020-02-24 |
| Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces | ✓ Link | 16.5 | Snips | 2018-05-25 |
| Semi-Supervised Speech Recognition via Local Prior Matching | ✓ Link | 20.84 | Local Prior Matching (Large Model) | 2020-02-24 |