| Qwen3.5-Omni Technical Report | | 82.2 | Qwen3.5-Omni-Plus | 2026-04-21 |
| Qwen3.5-Omni Technical Report | | 81.1 | Gemini-3.1 Pro | 2026-04-21 |
| Qwen3.5-Omni Technical Report | | 80.4 | Qwen3.5-Omni-Flash | 2026-04-21 |
| Announcing Amazon Nova 2 foundation models now available in Amazon Bedrock | | 77.80 | Nova 2 Omni | 2026-01-01 |
| Audio-Thinker | | 77.70 | Audio-Thinker (8.4B) | 2025-08-11 |
| Step-Audio-2 Technical Report | | 77.58 | Step-Audio-2 | 2025-07-22 |
| Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM? | | 77.00 | Omni-R1 (FT on Qwen2.5-Omni, 8.2B) | 2025-05-14 |
| MiMo-Audio: Audio Language Models are Few-Shot Learners | ✓ Link | 74.70 | MiMo-Audio (7B) | 2025-09-01 |
| Audio Flamingo 3 | | 73.30 | Audio Flamingo 3 (8.2B) | 2025-07-10 |
| Step-Audio-2 Technical Report | | 72.73 | Step-Audio-2-mini (8.3B) | 2025-07-22 |
| Start building with Gemini 2.5 Flash | | 71.80 | Gemini 2.5 Flash | 2025-06-17 |
| Gemini 2.5: Our newest Gemini model with thinking | | 71.60 | Gemini 2.5 Pro | 2025-03-25 |
| Qwen2.5-Omni Technical Report | | 71.50 | Qwen2.5-Omni (8.2B) | 2025-03-26 |
| Google introduces Gemini 2.0: A new AI model for the agentic era | | 70.50 | Gemini 2.0 Flash | 2024-12-11 |
| Kimi-Audio Technical Report | | 68.20 | Kimi-Audio (8.2B) | 2025-04-25 |
| Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models | ✓ Link | 67.70 | Audio Reasoner (8.2B) | 2025-02-01 |
| Gemini 2.5 Flash-Lite is now stable and generally available | | 66.20 | Gemini 2.5 Flash Lite | 2025-07-22 |
| DeSTA2.5-Audio | | 66.00 | DeSTA2.5-Audio (8B) | 2025-07-03 |
| Phi-4 Technical Report | ✓ Link | 65.70 | Phi-4-multimodal (5.5B) | 2025-02-27 |
| Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding | ✓ Link | 64.59 | Audio Flamingo 2 Reasoning (FT on AF2, 3B) | 2025-08-01 |
| GPT-4o System Card | | 62.50 | GPT-4o Audio | 2024-10-01 |
| Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities | ✓ Link | 62.40 | Audio Flamingo 2 (3B) | 2025-02-06 |
| Qwen2-Audio Technical Report | | 59.60 | Qwen2-Audio-Instruct (7B) | 2024-07-15 |
| Announcing Gemma 3n preview: powerful, efficient, mobile-first AI | ✓ Link | 58.00 | Gemma 3n (4B) | 2025-05-20 |
| GPT-4o System Card | | 53.00 | GPT-4o mini Audio | 2024-10-01 |
| Announcing Gemma 3n preview: powerful, efficient, mobile-first AI | ✓ Link | 51.69 | Gemma 3n (2B) | 2025-05-20 |
| MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response | ✓ Link | 38.20 | MusiLingo (7B) | 2023-09-15 |
| M2UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models | ✓ Link | 37.90 | M2UGen (7B) | 2023-08-18 |
| SALMONN: Towards Generic Hearing Abilities for LLMs | | 34.90 | SALMONN (13B) | 2023-10-20 |
| MU-LLaMA: Music Understanding LLaMA | | 27.60 | MuLLaMa (7B) | 2023-08-22 |
| GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities | ✓ Link | 22.83 | GAMA-IT (7B) | 2024-06-17 |
| GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities | ✓ Link | 20.82 | GAMA (7B) | 2024-06-17 |
| LTU: Listen, Think, and Understand | | 17.44 | LTU (7B) | 2023-09-01 |
| Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities | ✓ Link | 16.60 | Audio Flamingo Chat (1B) | 2024-02-02 |