| OmniVec: Learning robust representations with cross modal sharing | | 0.552 | OmniVec (audio+visual) | 2023-11-09 |
| OmniVec: Learning robust representations with cross modal sharing | | 0.548 | OmniVec (audio-only) | 2023-11-09 |
| EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning | ✓ Link | 0.546 | EquiAV (audio-visual) | 2024-03-14 |
| Contrastive Audio-Visual Masked Autoencoder | ✓ Link | 0.512 | CAV-MAE (Audio-Visual) | 2022-10-02 |
| BEATs: Audio Pre-Training with Acoustic Tokenizers | ✓ Link | 0.506 | BEATs (10 models) | 2022-12-18 |
| CED: Consistent ensemble distillation for audio tagging | ✓ Link | 0.500 | CED-Base | 2023-08-22 |
| Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation | ✓ Link | 0.498 | mn40_as (Ensemble) | 2022-11-09 |
| Efficient Training of Audio Transformers with Patchout | ✓ Link | 0.496 | PaSST | 2021-10-11 |
| Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models | ✓ Link | 0.490 | DyMN-L (Audio-Only, Single) | 2023-10-24 |
| HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection | ✓ Link | 0.487 | HTS-AT (Ensemble) | 2022-02-02 |
| AST: Audio Spectrogram Transformer | ✓ Link | 0.485 | Audio Spectrogram Transformer | 2021-04-05 |
| Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation | ✓ Link | 0.483 | mn40_as (Single) | 2022-11-09 |
| PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation | ✓ Link | 0.474 | PSLA | 2021-02-02 |
| Zero-shot Audio Source Separation through Query-based Learning from Weakly-labeled Data | ✓ Link | 0.467 | ST-SED | 2021-12-15 |
| Contrastive Audio-Visual Masked Autoencoder | ✓ Link | 0.466 | CAV-MAE (Audio-Only) | 2022-10-02 |
| ERANNs: Efficient Residual Audio Neural Networks for Audio Pattern Recognition | | 0.450 | ERANN-1-6 | 2021-06-03 |
| Perceiver IO: A General Architecture for Structured Inputs & Outputs | ✓ Link | 0.450 | Perceiver IO (mel-spectrogram + video) | 2021-07-30 |
| Perceiver: General Perception with Iterative Attention | ✓ Link | 0.440 | Perceiver (mel spectrogram + video - tuned) | 2021-03-04 |
| PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition | ✓ Link | 0.431 | CNN14 | 2020-08-23 |
| Audio Mamba: Bidirectional State Space Model for Audio Representation Learning | ✓ Link | 0.324 | AuM-B/16 | 2024-06-05 |
| MiDashengLM: Efficient Audio Understanding with General Audio Captions | ✓ Link | 0.089 | MiDashengLM-7B | 2025-08-06 |