OpenCodePapers

speech-recognition-on-gigaspeech-test

Speech Recognition
Dataset Link
Results over time
Click legend items to toggle metrics. Hover points for model names.
Leaderboard
PaperCodeWord Error Rate (WER)ModelNameReleaseDate
6.75OpenMOSS-Team/MOSS-Transcribe-preview-2B2026-06-26
Data-Efficient On-Policy Distillation for Automatic Speech Recognition✓ Link6.98AutoArk-AI/ARK-ASR-3B2026-06-22
Boson AI Launches Higgs STT 3 Speech-to-Text Model✓ Link7.13bosonai/higgs-audio-v3-8b-stt-v22026-04-27
Qwen3-ASR Technical Report✓ Link7.25Qwen/Qwen3-ASR-1.7B2026-01-28
Azure Speech at Build 2026: Powering Voice Agents with Real-Time and Life-like Experiences7.34microsoft/azure-speech-06-20262026-06-04
Data-Efficient On-Policy Distillation for Automatic Speech Recognition✓ Link7.42AutoArk-AI/ARK-ASR-0.6B2026-05-25
Introducing Resonant-1 and Resonant-1-flash7.44reson8/resonant-12026-04-08
Introducing Resonant-1 and Resonant-1-flash7.47reson8/resonant-1-flash2026-04-08
Boson AI Launches Higgs STT 3 Speech-to-Text Model✓ Link7.49bosonai/higgs-audio-v3-stt2026-03-18
Introducing Universal-3 Pro: A new class of speech language model optimized for Voice AI7.6assemblyai/universal-3-pro2026-02-03
Qwen3-ASR Technical Report✓ Link7.65Qwen/Qwen3-ASR-0.6B2026-01-28
✓ Link7.88nvidia/canary-qwen-2.5b2025-07-17
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs7.91microsoft/Phi-4-multimodal-instruct2025-02-24
Introducing Cohere-transcribe: state-of-the-art speech recognition7.93CohereLabs/cohere-transcribe-03-20262026-03-26
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link7.98nvidia/parakeet-tdt-1.1b2024-01-25
Ursa 2: Elevating speech recognition across 50+ languages7.98speechmatics/enhanced2024-10-11
Introducing Zoom AI Services8.04zoom/scribe_v12026-03-10
VibeVoice-ASR Technical Report✓ Link8.05microsoft/VibeVoice-ASR-HF2026-03-02
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST✓ Link8.07nvidia/parakeet-tdt-0.6b-v32025-08-14
Avalon: ASR for Human–AI Interaction8.07aquavoice/avalon-v1-en2025-08-22
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link8.14usefulsensors/moonshine-streaming-medium2026-02-12
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link8.23nvidia/parakeet-tdt-0.6b-v22025-05-01
GLM-ASR-Nano: A robust, open-source speech recognition model✓ Link8.23zai-org/GLM-ASR-Nano-25122025-12-09
✓ Link8.23ibm-granite/granite-speech-4.1-2b2026-04-29
Robust Knowledge Distillation via Large-Scale Pseudo Labelling✓ Link8.26distil-whisper/distil-large-v3.52024-12-05
Training and Inference Efficiency of Encoder-Decoder Speech Models✓ Link8.26nvidia/canary-1b-flash2025-03-07
Pulse STT: Fast, Accurate Speech-to-Text8.26smallestai/pulse2026-01-28
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling✓ Link8.28kyutai/stt-2.6b-en2025-06-06
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link8.31nvidia/parakeet-rnnt-1.1b2023-12-27
Voxtral✓ Link8.37mistralai/Voxtral-Small-24B-25072025-07-15
Robust Speech Recognition via Large-Scale Weak Supervision✓ Link8.4openai/whisper-large-v32023-11-06
✓ Link8.51ibm-granite/granite-4.0-1b-speech2026-03-06
Robust Speech Recognition via Large-Scale Weak Supervision✓ Link8.52openai/whisper-large-v3-turbo2024-10-01
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link8.53efficient-speech/lite-whisper-large-v3-acc2025-02-26
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link8.55nvidia/parakeet-rnnt-0.6b2023-12-28
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data✓ Link8.55nvidia/canary-1b2024-02-07
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link8.64efficient-speech/lite-whisper-large-v3-turbo-acc2025-02-26
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities✓ Link8.65ibm-granite/granite-speech-3.3-8b2025-06-19
NLE: Non-autoregressive LLM-based ASR by Transcript Editing✓ Link8.67ibm-granite/granite-speech-4.1-2b-nar2026-04-01
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link8.69efficient-speech/lite-whisper-large-v32025-02-26
Voxtral✓ Link8.75mistralai/Voxtral-Mini-3B-25072025-07-01
Voxtral Realtime✓ Link8.86mistralai/Voxtral-Mini-4B-Realtime-26022026-02-11
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link8.88nvidia/parakeet-ctc-1.1b2023-12-28
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations✓ Link8.93nvidia/parakeet-tdt_ctc-110m2024-09-17
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions✓ Link8.96nyrahealth/CrisperWhisper2024-08-29
Training and Inference Efficiency of Encoder-Decoder Speech Models✓ Link8.96nvidia/canary-180m-flash2025-03-11
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link9.03nvidia/parakeet-ctc-0.6b2023-12-28
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link9.08usefulsensors/moonshine-streaming-small2026-02-12
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation✓ Link9.09efficient-speech/lite-whisper-large-v3-fast2025-02-26
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities✓ Link9.1ibm-granite/granite-speech-3.3-2b2025-06-19
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST✓ Link9.2nvidia/canary-1b-v22025-08-14
Zipformer: A faster and better encoder for automatic speech recognition✓ Link9.31soundsgoodai/Zipformer-transducer-XL-290M2026-05-12
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning✓ Link9.37espnet/owsm_ctc_v4_1B2025-01-16
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages✓ Link9.5facebook/omniASR-LLM-7B-v22025-12-12
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link10.03Zipformer+pruned transducer w/ CR-CTC (no external language model)2024-10-07
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link10.07Zipformer+CR-CTC/AED (no external language model)2024-10-07
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link10.2Zipformer+pruned transducer (no external language model)2024-10-07
CR-CTC: Consistency regularization on CTC for improved speech recognition✓ Link10.28Zipformer+CR-CTC (no external language model)2024-10-07
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models✓ Link10.29espnet/owsm_ctc_v3.2_ft_1B2024-09-24
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link10.31nvidia/stt_en_conformer_ctc_large2022-04-09
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification✓ Link10.44espnet/owsm_ctc_v3.1_1B2024-02-23
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use✓ Link10.48speechbrain/asr-conformer-largescaleasr2025-02-06
Moonshine: Speech Recognition for Live Transcription and Voice Commands✓ Link10.69usefulsensors/moonshine-base2024-10-21
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link10.76nvidia/stt_en_fastconformer_transducer_large2023-06-08
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio✓ Link10.80Conformer/Transformer-AED2021-06-13
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition✓ Link10.98nvidia/stt_en_fastconformer_ctc_large2023-06-08
Niagara-38m Sets a New Benchmark for Edge Speech Recognition11.41abr-ai/niagara-38m-batch.en2026-04-15
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link12.43nvidia/stt_en_conformer_transducer_small2022-06-01
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications✓ Link12.53usefulsensors/moonshine-streaming-tiny2026-02-12
Moonshine: Speech Recognition for Live Transcription and Voice Commands✓ Link12.72usefulsensors/moonshine-tiny2024-10-21
Conformer: Convolution-augmented Transformer for Speech Recognition✓ Link13.32nvidia/stt_en_conformer_ctc_small2023-06-12
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages✓ Link13.36facebook/omniASR-CTC-7B-v22025-12-12
Niagara-38m Sets a New Benchmark for Edge Speech Recognition14.21abr-ai/niagara-19m-batch.en2026-04-15
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link15.01facebook/wav2vec2-large-960h-lv60-self2020-06-20
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units✓ Link16.01facebook/hubert-xlarge-ls960-ft2022-03-02
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language✓ Link16.01facebook/data2vec-audio-large-960h2022-04-02
fairseq S2T: Fast Speech-to-Text Modeling with fairseq✓ Link16.19facebook/wav2vec2-conformer-rel-pos-large-960h-ft2022-04-18
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units✓ Link16.28facebook/hubert-large-ls960-ft2022-03-02
fairseq S2T: Fast Speech-to-Text Modeling with fairseq✓ Link16.36facebook/wav2vec2-conformer-rope-large-960h-ft2022-04-18
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training✓ Link16.52facebook/wav2vec2-large-robust-ft-libri-960h2021-04-02
SpeechBrain: A General-Purpose Speech Toolkit✓ Link16.72speechbrain/asr-wav2vec2-librispeech2022-06-05
Scaling Speech Technology to 1,000+ Languages✓ Link17.47facebook/mms-1b-all2023-05-27
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link19.37facebook/wav2vec2-large-960h2020-06-20
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language✓ Link21.7facebook/data2vec-audio-base-960h2022-03-02
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations✓ Link22.58facebook/wav2vec2-base-960h2022-03-02