| DINOv3 | ✓ Link | 66.1 | | | | | | 7100 | DINOv3 7B/16 (Plain-DETR, frozen backbone, TTA) | 2025-08-13 |
| Perception Encoder: The best visual embeddings are not at the output of the network | ✓ Link | 66.0 | | | | | | 1900 | PE_spatial (DETA) | 2025-04-17 |
| DETRs with Collaborative Hybrid Assignments Training | ✓ Link | 65.9 | | | | | | 314 | Co-DETR | 2022-11-22 |
| DINOv3 | ✓ Link | 65.6 | | | | | | 7100 | DINOv3 7B/16 (Plain-DETR, frozen backbone, no TTA) | 2025-08-13 |
| InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions | ✓ Link | 65.0 | | | | | | | InternImage-H | 2022-11-10 |
| DETRs with Collaborative Hybrid Assignments Training | ✓ Link | 64.7 | | | | | | 218 | Co-DETR (Swin-L) | 2022-11-22 |
| A Strong and Reproducible Object Detector with Only Public Datasets | ✓ Link | 64.6 | 81.5 | 71.4 | 50.4 | 68.5 | 78.5 | 689 | Focal-Stable-DINO (Focal-Huge, no TTA) | 2023-04-25 |
| EVA: Exploring the Limits of Masked Visual Representation Learning at Scale | ✓ Link | 64.5 | | | | | | | EVA (Cascade Mask R-CNN, TTA) | 2022-11-14 |
| ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions | ✓ Link | 64.3 | | | | | | 363 | ViT-CoMer | 2024-03-12 |
| Focal Modulation Networks | ✓ Link | 64.2 | | | | | | | FocalNet-H (DINO) | 2022-03-22 |
| InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions | ✓ Link | 64.2 | | | | | | | InternImage-XL | 2022-11-10 |
| CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection | | 64.1 | | | | | | | CP-DETR-L Swin-L(Fine tuning separately in COCO) | 2024-12-13 |
| Reversible Column Networks | ✓ Link | 63.8 | | | | | | | RevCol-H(DINO) | 2022-12-22 |
| DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection | ✓ Link | 63.2 | | | | | | | DINO (Swin-L) | 2022-03-07 |
| Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection | ✓ Link | 63.0 | | | | | | | Grounding DINO L (1.5x image size) | 2023-03-09 |
| Swin Transformer V2: Scaling Up Capacity and Resolution | ✓ Link | 62.5 | | | | | | | SwinV2-G (HTC++) | 2021-11-18 |
| Florence: A New Foundation Model for Computer Vision | | 62 | | | | | | | Florence-CoSwin-H | 2021-11-22 |
| General Object Foundation Model for Images and Videos at Scale | ✓ Link | 62.0 | | | | | | | GLEE-Pro | 2023-12-14 |
| Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration | ✓ Link | 61.8 | | | | | | | Fast-iTPN-B (DINO, CLIP-distilled pre-training + Objects365 detection pre-training) | 2022-11-23 |
| Mr. DETR: Instructive Multi-Route Training for Detection Transformers | ✓ Link | 61.8 | 79.0 | 67.6 | 47.7 | 65.6 | 75.7 | | Mr. DETR-Align (Swin-L, Objects365 pre-training, 5-scale; official repo) | 2024-12-13 |
| Exploring Plain Vision Transformer Backbones for Object Detection | ✓ Link | 61.3 | | | | | | | ViTDet, ViT-H Cascade (multiscale) | 2022-03-30 |
| Grounded Language-Image Pre-training | ✓ Link | 60.8 | | | | | | | GLIP (Swin-L, multi-scale) | 2021-12-07 |
| End-to-End Semi-Supervised Object Detection with Soft Teacher | ✓ Link | 60.7 | | | | | | | Soft Teacher + Swin-L (HTC++, multi-scale) | 2021-06-16 |
| Universal Instance Perception as Object Discovery and Retrieval | ✓ Link | 60.6 | 77.5 | 66.7 | 45.1 | 64.8 | 75.3 | | UNINEXT-H | 2023-03-12 |
| Vision Transformer Adapter for Dense Predictions | ✓ Link | 60.5 | | | | | | | ViT-Adapter-L (HTC++, BEiTv2 pretrain, multi-scale) | 2022-05-17 |
| Exploring Plain Vision Transformer Backbones for Object Detection | ✓ Link | 60.4 | | | | | | | ViTDet, ViT-H Cascade | 2022-03-30 |
| General Object Foundation Model for Images and Videos at Scale | ✓ Link | 60.4 | | | | | | | GLEE-Plus | 2023-12-14 |
| Dynamic Head: Unifying Object Detection Heads with Attentions | ✓ Link | 60.3 | | | | | | | DyHead (Swin-L, multi scale, self-training) | 2021-06-15 |
| Vision Transformer Adapter for Dense Predictions | ✓ Link | 60.2 | | | | | | | ViT-Adapter-L (HTC++, BEiT pretrain, multi-scale) | 2022-05-17 |
| End-to-End Semi-Supervised Object Detection with Soft Teacher | ✓ Link | 60.1 | | | | | | | Soft Teacher+Swin-L(HTC++, single scale) | 2021-06-16 |
| Parameter-Inverted Image Pyramid Networks | ✓ Link | 60.0 | 79.0 | 65.4 | | | | | PIIP-H6B (DINO) | 2024-06-06 |
| SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation | ✓ Link | 59.8 | | | 42.2 | 63.8 | 74.9 | | SimPLR (ViT-H, MAE IN-1K pre-training, 1024px) | 2023-10-09 |
| CBNet: A Composite Backbone Network Architecture for Object Detection | ✓ Link | 59.6 | | | | | | | CBNetV2 (Dual-Swin-L HTC, multi-scale) | 2021-07-01 |
| Could Giant Pretrained Image Models Extract Universal Representations? | | 59.3 | | | | | | | Frozen Backbone, SwinV2-G-ext22K (HTC) | 2022-11-03 |
| HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions | ✓ Link | 59.2 | | | | | | | HorNet-L | 2022-07-28 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 59.2 | | | | | | | MOAT-3 (IN-22K pretraining, single-scale) | 2022-10-04 |
| CBNet: A Composite Backbone Network Architecture for Object Detection | ✓ Link | 59.1 | | | | | | | CBNetV2 (Dual-Swin-L HTC, multi-scale) | 2021-07-01 |
| Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration | ✓ Link | 58.8 | | | | | | | Fast-iTPN-L (DINO 12ep, CLIP-distilled pre-training) | 2022-11-23 |
| ScaleDet: A Scalable Multi-Dataset Object Detector | | 58.8 | | | | | | | ScaleDet-B (Swin-B, multi-dataset training) | 2023-06-08 |
| Focal Self-attention for Local-Global Interactions in Vision Transformers | ✓ Link | 58.7 | | | | | | | Focal-L (DyHead, multi-scale) | 2021-07-01 |
| MViTv2: Improved Multiscale Vision Transformers for Classification and Detection | ✓ Link | 58.7 | | | | | | | MViTv2-L (Cascade Mask R-CNN, multi-scale, IN21k pre-train) | 2021-12-02 |
| SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation | ✓ Link | 58.7 | | | 40.4 | 63.2 | 74.8 | | SimPLR (ViT-L, BEiTv2 IN-21K pre-training, 1024px) | 2023-10-09 |
| Mr. DETR: Instructive Multi-Route Training for Detection Transformers | ✓ Link | 58.7 | 76.5 | 64.0 | 42.2 | 62.9 | 75.2 | | Mr. DETR++ (Swin-L, 12ep) | 2024-12-13 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 58.5 | | | | | | | MOAT-2 (IN-22K pretraining, single-scale) | 2022-10-04 |
| SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation | ✓ Link | 58.5 | | | 42.2 | 62.5 | 73.4 | | SimPLR (ViT-L, MAE IN-1K pre-training, 1024px) | 2023-10-09 |
| Dynamic Head: Unifying Object Detection Heads with Attentions | ✓ Link | 58.4 | | | | | | | DyHead (Swin-L, multi scale) | 2021-06-15 |
| Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration | ✓ Link | 58.4 | | | | | | | Fast-iTPN-B (DINO 12ep, CLIP-distilled pre-training) | 2022-11-23 |
| Mr. DETR: Instructive Multi-Route Training for Detection Transformers | ✓ Link | 58.4 | 76.3 | 63.9 | 40.8 | 62.8 | 75.3 | | Mr. DETR (Swin-L, 12ep) | 2024-12-13 |
| FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation | ✓ Link | 58.3 | 77.1 | 64.0 | 41.7 | 62.4 | 73.9 | | H-Deformable-DETR + FeatAug-Flip (Swin-L, 24ep, 300 predictions) | 2023-03-02 |
| MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism | ✓ Link | 58.2 | 76.5 | 63.4 | 42.5 | 62.8 | 74.6 | | MI-DETR (Swin-L, 12ep) | 2025-03-03 |
| Relation DETR: Exploring Explicit Position Relation Prior for Object Detection | ✓ Link | 58.1 | 76.4 | 63.5 | 41.8 | 63.0 | 73.5 | | Relation-DETR (Swin-L, 24ep; official repo) | 2024-07-16 |
| LP-DETR: Layer-wise Progressive Relations for Object Detection | | 58.1 | 76.4 | 63.2 | 41.0 | 62.2 | 74.7 | | LP-DETR (Swin-L, 12ep) | 2025-02-07 |
| MDS-DETR: DETR with Masked Duplicate Suppressor | ✓ Link | 58.1 | 76.2 | 63.6 | 42.6 | 62.7 | 74.7 | | MDS-DETR (Swin-L) | 2026-05-22 |
| Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | ✓ Link | 58 | | | | | | | Swin-L (HTC++, multi scale) | 2021-03-25 |
| Fractional Correspondence Framework in Detection Transformer | | 57.9 | 76.1 | 63.6 | 41.5 | 62.3 | 74.9 | | DINO + RTP (Swin-L, 12ep) | 2025-03-06 |
| Relation DETR: Exploring Explicit Position Relation Prior for Object Detection | ✓ Link | 57.8 | 76.1 | 62.9 | 41.2 | 62.1 | 74.4 | | Relation-DETR (Swin-L, 12ep) | 2024-07-16 |
| DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model | | 57.8 | | | 41.5 | 61.2 | 73.9 | | DI-MaskDINO (Swin-L, 50ep) | 2024-10-22 |
| PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection | | 57.8 | 76.2 | 62.6 | 41.4 | 62.2 | 74.2 | | PaQ-DINO (Swin-L, 12ep) | 2026-03-06 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 57.7 | | | | | | | MOAT-1 (IN-1K pretraining, single-scale) | 2022-10-04 |
| FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation | ✓ Link | 57.6 | 76.7 | 63.1 | 41.1 | 61.5 | 73.8 | | Deformable-DETR w/ tricks + FeatAug-FC (Swin-L, 24ep, 300 predictions) | 2023-03-02 |
| Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models | ✓ Link | 57.6 | 75.8 | 63.2 | 41.4 | 61.7 | 74.3 | | Frozen-DETR (Co-DINO-4scale, Swin-B IN-22K, frozen foundation model, 12ep) | 2024-10-25 |
| Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders | ✓ Link | 57.6 | 76.4 | 63.2 | 40.8 | 61.8 | 73.6 | | Dual-R-DETR (Swin-L, 12ep) | 2025-12-15 |
| Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality | ✓ Link | 57.4 | | | | | | | UM-MAE(HTC++, Swin-L, IN1K) | 2022-05-20 |
| Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement | ✓ Link | 57.3 | 75.5 | 62.3 | 40.9 | 61.8 | 74.5 | | Salience DETR (FocalNet-L, 12ep; official repo) | 2024-03-24 |
| YOLOv6 v3.0: A Full-Scale Reloading | ✓ Link | 57.2 | 74.5 | | | | | | YOLOv6-L6(46 fps, 1280, V100) | 2023-01-13 |
| Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | ✓ Link | 57.1 | | | | | | | Swin-L (HTC++, single scale) | 2021-03-25 |
| TransNeXt: Robust Foveal Visual Perception for Vision Transformers | ✓ Link | 57.1 | | | | | | | TransNeXt-Base (IN-1K pretrain, DINO 1x) | 2023-11-28 |
| Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | ✓ Link | 57.0 | | | | | | | Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale) | 2020-12-13 |
| TransNeXt: Robust Foveal Visual Perception for Vision Transformers | ✓ Link | 56.6 | | | | | | | TransNeXt-Small (IN-1K pretrain, DINO 1x) | 2023-11-28 |
| UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale | ✓ Link | 56.6 | 75.6 | 61.8 | | | | 254.8 | UniConvNet-L (IN-22K, Cascade Mask R-CNN 3x) | 2025-08-12 |
| Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement | ✓ Link | 56.5 | 75.0 | 61.5 | 40.2 | 61.2 | 72.8 | | Salience DETR (Swin-L, 12ep; official repo) | 2024-03-24 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 56.4 | | | | | | 443 | UniRepLKNet-XL (IN-22K, Cascade Mask R-CNN 3x) | 2023-11-27 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 56.2 | 75.0 | 61.2 | | | | 238 | MogaNet-XL (IN-22K, Cascade Mask R-CNN 3x) | 2022-11-07 |
| Instances as Queries | ✓ Link | 56.1 | | | | | | | QueryInst (Swin-L) | 2021-05-05 |
| MViTv2: Improved Multiscale Vision Transformers for Classification and Detection | ✓ Link | 56.1 | | | | | | | MViTv2-H (Cascade Mask R-CNN, single-scale, IN21k pre-train) | 2021-12-02 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 55.9 | | | | | | | MOAT-0 (IN-1K pretraining, single-scale) | 2022-10-04 |
| Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration | ✓ Link | 55.8 | | | | | | | Fast-iTPN-S (DINO 12ep, CLIP-distilled pre-training) | 2022-11-23 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 55.8 | | | | | | 276 | UniRepLKNet-L (IN-22K, Cascade Mask R-CNN 3x) | 2023-11-27 |
| TransNeXt: Robust Foveal Visual Perception for Vision Transformers | ✓ Link | 55.7 | | | | | | | TransNeXt-Tiny (IN-1K pretrain, DINO 1x) | 2023-11-28 |
| Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models | ✓ Link | 55.7 | 73.9 | 61.3 | 38.4 | 58.8 | 72.3 | | Frozen-DETR (DDQ + HPR, ResNet-50, frozen foundation model, 12ep) | 2024-10-25 |
| UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale | ✓ Link | 55.7 | 74.4 | 60.4 | | | | 254.8 | UniConvNet-L (IN-22K, Cascade Mask R-CNN 1x) | 2025-08-12 |
| Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens | | 55.4 | 74.1 | 59.9 | | | | 114 | SECViT-B (Cascade Mask R-CNN 3x) | 2024-05-22 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 55.2 | | | | | | | tiny-MOAT-3 (IN-1K pretraining, single-scale) | 2022-10-04 |
| Understanding The Robustness in Vision Transformers | ✓ Link | 55.1 | | | | | | | FAN-L-Hybrid | 2022-04-26 |
| Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles | ✓ Link | 55 | | | | | | | Hiera-L | 2023-06-01 |
| General Object Foundation Model for Images and Videos at Scale | ✓ Link | 55.0 | | | | | | | GLEE-Lite | 2023-12-14 |
| DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection | ✓ Link | 55.0 | 72.5 | 60.0 | 38.4 | 58.9 | 69.9 | 45 | DS-Det (Strip-MLP-T, 36ep) | 2025-07-26 |
| DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection | ✓ Link | 54.9 | 72.4 | 59.9 | 38.3 | 58.6 | 69.6 | 49 | DS-Det (Swin-T, 36ep) | 2025-07-26 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 54.8 | | | | | | 155 | UniRepLKNet-B (IN-22K, Cascade Mask R-CNN 3x) | 2023-11-27 |
| Towards Sustainable Self-supervised Learning | ✓ Link | 54.6 | | | | | | | TEC(VIT-B, Mask-RCNN) | 2022-10-20 |
| Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | ✓ Link | 54.5 | | | | | | | Cascade Eff-B7 NAS-FPN (1280) | 2020-12-13 |
| Context Autoencoder for Self-Supervised Representation Learning | ✓ Link | 54.5 | 75.2 | 60.1 | | | | | CAE* (ViT-L, 1600ep, Mask R-CNN 1x) | 2022-02-07 |
| RMT: Retentive Networks Meet Vision Transformers | ✓ Link | 54.5 | 72.8 | 59.0 | | | | 111 | RMT-B (Cascade Mask R-CNN 3x) | 2023-09-20 |
| Rotary Position Embedding for Vision Transformer | ✓ Link | 54.5 | | | | | | | Swin-B + RoPE-Mixed (DINO, 12ep) | 2024-03-20 |
| Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection | | 54.5 | 72.5 | 59.2 | 37.8 | 58.7 | 71.3 | | BS-O2G (DEIM, ResNet-50, 60ep) | 2026-08-06 |
| EfficientDet: Scalable and Efficient Object Detection | ✓ Link | 54.4 | | | | | | | EfficientDet-D7x (single-scale) | 2019-11-20 |
| MViTv2: Improved Multiscale Vision Transformers for Classification and Detection | ✓ Link | 54.3 | | | | | | | MViTv2-L (Cascade Mask R-CNN, single-scale) | 2021-12-02 |
| MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers | ✓ Link | 54.3 | | | | | | 319 | MixMAE (Swin-L, 600ep, Mask R-CNN) | 2022-05-26 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 54.3 | | | | | | 113 | UniRepLKNet-S (IN-22K, Cascade Mask R-CNN 3x) | 2023-11-27 |
| Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models | ✓ Link | 54.3 | 72.9 | 59.2 | 36.6 | 58.0 | 72.1 | | Frozen-DETR (Co-DINO-4scale, ResNet-50, frozen foundation model, 24ep) | 2024-10-25 |
| Rethinking Pre-training and Self-training | ✓ Link | 54.2 | | | | | | | SpineNet-190 (1280, with Self-training on OpenImages, single-scale) | 2020-06-11 |
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | | 53.9 | | | | | | 154 | OverLoCK-B (Cascade Mask R-CNN 3x) | 2025-02-27 |
| Spectral-Adaptive Modulation Networks for Visual Perception | ✓ Link | 53.8 | 72.9 | 58.4 | | | | | SPANetV2-S36-hybrid (Cascade Mask R-CNN 3x) | 2025-03-31 |
| Simple Training Strategies and Model Scaling for Object Detection | ✓ Link | 53.6 | | | 34.5 | 56.7 | 70.6 | | Cascade RCNN-RS (SpineNet-143L, single scale) | 2021-06-30 |
| Exploring Target Representations for Masked Autoencoders | ✓ Link | 53.6 | | | | | | | dBOT (ViT-B, CLIP teacher, Mask R-CNN) | 2022-09-08 |
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | | 53.6 | | | | | | 114 | OverLoCK-S (Cascade Mask R-CNN 3x) | 2025-02-27 |
| Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection | | 53.5 | 71.3 | 58.1 | 36.9 | 58.0 | 70.5 | | BS-O2G (DEIM, ResNet-50, 24ep) | 2026-08-06 |
| EfficientDet: Scalable and Efficient Object Detection | ✓ Link | 53.4 | | | | | | | EfficientDet-D7 (1536) | 2019-11-20 |
| MaxViT: Multi-Axis Vision Transformer | ✓ Link | 53.4 | 72.9 | 58.1 | | | | 157 | MaxViT-B (Cascade Mask R-CNN) | 2022-04-04 |
| Masked Autoencoders Are Scalable Vision Learners | ✓ Link | 53.3 | | | | | | | MAE (ViT-L, Mask R-CNN) | 2021-11-11 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 53.3 | 71.8 | 57.8 | | | | 140 | MogaNet-L (Cascade Mask R-CNN 3x) | 2022-11-07 |
| RMT: Retentive Networks Meet Vision Transformers | ✓ Link | 53.2 | 72.0 | 57.8 | | | | 83 | RMT-S (Cascade Mask R-CNN 3x) | 2023-09-20 |
| Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models | ✓ Link | 53.2 | 71.8 | 58.0 | 35.1 | 56.5 | 70.6 | | Frozen-DETR (DINO-4scale, ResNet-50, frozen foundation model, 24ep) | 2024-10-25 |
| Simple Training Strategies and Model Scaling for Object Detection | ✓ Link | 53.1 | | | 33.9 | 56.2 | 70.3 | | Cascade RCNN-RS (ResNet-200, single scale) | 2021-06-30 |
| MaxViT: Multi-Axis Vision Transformer | ✓ Link | 53.1 | 72.5 | 58.1 | | | | 107 | MaxViT-S (Cascade Mask R-CNN) | 2022-04-04 |
| PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution | | 53.1 | | | | | | 147 | PeLK-B-101 (Cascade Mask R-CNN 3x) | 2024-03-12 |
| Spectral-Adaptive Modulation Networks for Visual Perception | ✓ Link | 53.1 | 72.2 | 57.6 | | | | | SPANetV2-S36-pure (Cascade Mask R-CNN 3x) | 2025-03-31 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 53.0 | | | | | | | tiny-MOAT-2 (IN-1K pretraining, single-scale) | 2022-10-04 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 53.0 | | | | | | 113 | UniRepLKNet-S (Cascade Mask R-CNN 3x) | 2023-11-27 |
| PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution | | 52.9 | | | | | | 147 | PeLK-B (Cascade Mask R-CNN 3x) | 2024-03-12 |
| Rotary Position Embedding for Vision Transformer | ✓ Link | 52.9 | | | | | | | ViT-L + RoPE-Mixed (DINO-ViTDet, 12ep) | 2024-03-20 |
| DeepMIM: Deep Supervision for Masked Image Modeling | ✓ Link | 52.8 | | | | | | | DeepMIM-MAE-CLIP (ViT-B, Mask R-CNN) | 2023-03-15 |
| Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens | | 52.8 | 73.6 | 57.7 | | | | 75 | SECViT-B (Mask R-CNN 3x) | 2024-05-22 |
| MambaVision: A Hybrid Mamba-Transformer Vision Backbone | ✓ Link | 52.8 | 71.3 | 57.2 | | | | 145 | MambaVision-B (Cascade Mask R-CNN 3x) | 2024-07-10 |
| MViTv2: Improved Multiscale Vision Transformers for Classification and Detection | ✓ Link | 52.7 | | | | | | | MViT-L (Mask R-CNN, single-scale, IN21k pre-train) | 2021-12-02 |
| MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers | ✓ Link | 52.7 | | | | | | 110 | MixMAE (Swin-B/W14, 600ep, Mask R-CNN) | 2022-05-26 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 52.6 | 72.0 | 57.3 | | | | 101 | MogaNet-B (Cascade Mask R-CNN 3x) | 2022-11-07 |
| HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs | | 52.6 | 71.3 | 57.1 | | | | 88.9 | HIRI-ViT-S (Cascade Mask R-CNN 3x, 1600px input) | 2024-03-18 |
| PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection | | 52.6 | 69.7 | 56.9 | 35.7 | 56.4 | 67.0 | | PaQ-DINO (ResNet-50, 24ep) | 2026-03-06 |
| Active Token Mixer | ✓ Link | 52.5 | 71.6 | 57.0 | | | | | ATMNet-B (Cascade Mask R-CNN 3x, MS) | 2022-03-11 |
| LP-DETR: Layer-wise Progressive Relations for Object Detection | | 52.5 | 70.0 | 57.2 | 36.2 | 56.3 | 67.1 | | LP-DETR (ResNet-50, 24ep) | 2025-02-07 |
| Language-aware Multiple Datasets Detection Pretraining for DETRs | | 52.4 | 70.3 | 57.5 | 35.7 | 55.9 | 66.4 | 50 | METR (ResNet-50, Objects365+OpenImages pre-training, 12ep) | 2023-04-07 |
| POA: Pre-training Once for Models of All Sizes | ✓ Link | 52.4 | | | | | | | POA ViT-B/16 (Cascade Mask R-CNN) | 2024-08-02 |
| PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection | | 52.4 | 69.8 | 57.4 | 36.1 | 55.6 | 66.2 | | PaQ-DETR (ResNet-50, hybrid matching, 12ep) | 2026-03-06 |
| MambaVision: A Hybrid Mamba-Transformer Vision Backbone | ✓ Link | 52.3 | 71.1 | 56.7 | | | | 108 | MambaVision-S (Cascade Mask R-CNN 3x) | 2024-07-10 |
| LP-DETR: Layer-wise Progressive Relations for Object Detection | | 52.3 | 69.6 | 56.8 | 35.8 | 55.9 | 66.6 | | LP-DETR (ResNet-50, 12ep) | 2025-02-07 |
| MDS-DETR: DETR with Masked Duplicate Suppressor | ✓ Link | 52.3 | 69.7 | 56.9 | 35.5 | 56.0 | 67.0 | | MDS-DETR (ResNet-50, 900 queries, 24ep) | 2026-05-22 |
| SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization | ✓ Link | 52.2 | | | | | | | Mask R-CNN (SpineNet-190, 1536x1536) | 2019-12-10 |
| RMT: Retentive Networks Meet Vision Transformers | ✓ Link | 52.2 | 72.9 | 57.0 | | | | 73 | RMT-B (Mask R-CNN 3x) | 2023-09-20 |
| PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution | | 52.2 | | | | | | 108 | PeLK-S (Cascade Mask R-CNN 3x) | 2024-03-12 |
| CoCAViT: Compact Vision Transformer with Robust Global Coordination | | 52.2 | 71.0 | 56.8 | | | | | CoCAViT-28M (Cascade Mask R-CNN 3x) | 2025-08-07 |
| MaxViT: Multi-Axis Vision Transformer | ✓ Link | 52.1 | 71.9 | 56.8 | | | | 69 | MaxViT-T (Cascade Mask R-CNN) | 2022-04-04 |
| Relation DETR: Exploring Explicit Position Relation Prior for Object Detection | ✓ Link | 52.1 | 69.7 | 56.6 | 36.1 | 56.0 | 66.5 | | Relation-DETR (ResNet-50, 24ep) | 2024-07-16 |
| Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens | | 52.0 | 73.5 | 57.3 | | | | 119 | SECViT-L (Mask R-CNN 1x) | 2024-05-22 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 51.9 | | | | | | | tiny-MOAT-1 (IN-1K pretraining, single-scale) | 2022-10-04 |
| HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs | | 51.9 | 71.9 | 56.6 | | | | 113.0 | HIRI-ViT-S (Sparse R-CNN 3x, 1600px input) | 2024-03-18 |
| DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model | | 51.9 | | | 36.3 | 54.7 | 66.7 | | DI-MaskDINO (ResNet-50, 50ep) | 2024-10-22 |
| PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection | | 51.9 | 69.1 | 56.3 | 35.1 | 56.0 | 66.6 | | PaQ-DINO (ResNet-50, 12ep) | 2026-03-06 |
| Global Context Networks | ✓ Link | 51.8 | 70.4 | 56.1 | | | | | GCNet (ResNeXt-101 + DCN + cascade + GC r4) | 2020-12-24 |
| UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition | ✓ Link | 51.8 | | | | | | 89 | UniRepLKNet-T (Cascade Mask R-CNN 3x) | 2023-11-27 |
| LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking | | 51.8 | 71.6 | 56.4 | 35.4 | 55.0 | 68.4 | 85 | AdaMixer + LORS (Swin-S, 3x) | 2024-03-07 |
| Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation | | 51.8 | 70.4 | 56.7 | 34.6 | 54.4 | 67.5 | | DINO + SACQ (Swin-T, 12ep) | 2024-05-06 |
| CoCAViT: Compact Vision Transformer with Robust Global Coordination | | 51.8 | 70.5 | 56.1 | | | | | CoCAViT-21M (Cascade Mask R-CNN 3x) | 2025-08-07 |
| Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss | ✓ Link | 51.7 | 69.0 | 56.3 | 35.5 | 55.0 | 66.1 | | Align-DETR (ResNet-50, 24ep) | 2023-04-15 |
| Relation DETR: Exploring Explicit Position Relation Prior for Object Detection | ✓ Link | 51.7 | 69.1 | 56.3 | 36.1 | 55.6 | 66.1 | | Relation-DETR (ResNet-50, 12ep) | 2024-07-16 |
| ELSA: Enhanced Local Self-Attention for Vision Transformer | ✓ Link | 51.6 | 70.5 | 56.0 | | | | | ELSA-S (Cascade Mask RCNN) | 2021-12-23 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 51.6 | 70.8 | 56.3 | | | | 83 | MogaNet-S (Cascade Mask R-CNN 3x) | 2022-11-07 |
| DeepMIM: Deep Supervision for Masked Image Modeling | ✓ Link | 51.6 | | | | | | | DeepMIM-MAE (ViT-B, Mask R-CNN) | 2023-03-15 |
| RMT: Retentive Networks Meet Vision Transformers | ✓ Link | 51.6 | 73.1 | 56.5 | | | | 114 | RMT-L (Mask R-CNN 1x) | 2023-09-20 |
| Focal Modulation Networks | ✓ Link | 51.5 | 70.3 | 56.0 | | | | | FocalNet-T (LRF, Cascade Mask R-CNN) | 2022-03-22 |
| SpectFormer: Frequency and Attention is what you need in a Vision Transformer | | 51.5 | 70.2 | 56.3 | | | | | SpectFormer-H-S (Cascade Mask R-CNN 3x) | 2023-04-13 |
| PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution | | 51.4 | | | | | | 86 | PeLK-T (Cascade Mask R-CNN 3x) | 2024-03-12 |
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | | 51.4 | | | | | | 114 | OverLoCK-B (Mask R-CNN 3x) | 2025-02-27 |
| DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection | ✓ Link | 51.3 | 69.1 | 56 | 34.5 | 54.2 | 65.8 | | DINO-5scale (24 epoch) | 2022-03-07 |
| DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection | ✓ Link | 51.3 | 68.4 | 55.9 | 35.5 | 54.8 | 65.7 | 48 | DS-Det (ResNet-50, 24ep) | 2025-07-26 |
| DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection | ✓ Link | 51.2 | 69 | 55.8 | 35 | 54.3 | 65.3 | | DINO-5scale (36 epoch) | 2022-03-07 |
| Rotary Position Embedding for Vision Transformer | ✓ Link | 51.2 | | | | | | | ViT-B + RoPE-Mixed (DINO-ViTDet, 12ep) | 2024-03-20 |
| UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale | ✓ Link | 51.2 | 72.2 | 56.1 | | | | 118.0 | UniConvNet-B (Mask R-CNN 3x) | 2025-08-12 |
| MDS-DETR: DETR with Masked Duplicate Suppressor | ✓ Link | 51.2 | 68.9 | 55.8 | 34.0 | 54.8 | 66.2 | | Hybrid-MDS-DETR (ResNet-50, 12ep) | 2026-05-22 |
| MambaVision: A Hybrid Mamba-Transformer Vision Backbone | ✓ Link | 51.1 | 70.0 | 55.6 | | | | 86 | MambaVision-T (Cascade Mask R-CNN 3x) | 2024-07-10 |
| MDS-DETR: DETR with Masked Duplicate Suppressor | ✓ Link | 51.1 | 68.8 | 55.5 | 34.0 | 55.0 | 66.0 | | MDS-DETR (ResNet-50, 900 queries, 12ep) | 2026-05-22 |
| ResNeSt: Split-Attention Networks | ✓ Link | 50.91 | | | | | | | ResNeSt-200 (Cascade R-CNN, DCN, tricks, 3x; official repo) | 2020-04-19 |
| RF-Next: Efficient Receptive Field Search for Convolutional Neural Networks | ✓ Link | 50.9 | 69.5 | 55.5 | 34.3 | 54.6 | 65.8 | | RF-ConvNeXt-T (Cascade Mask R-CNN) | 2022-06-14 |
| Learning Data Augmentation Strategies for Object Detection | ✓ Link | 50.7 | | | 34.2 | 55.5 | 64.5 | | NAS-FPN (AmoebaNet-D, learned aug) | 2019-06-26 |
| Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders | ✓ Link | 50.7 | 68.8 | 55.0 | 34.1 | 54.2 | 64.6 | | Dual-R-DETR (ResNet-50, 12ep) | 2025-12-15 |
| POA: Pre-training Once for Models of All Sizes | ✓ Link | 50.6 | | | | | | | POA ViT-S/16 (Cascade Mask R-CNN) | 2024-08-02 |
| Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer | | 50.6 | 71.3 | 55.3 | | | | 140 | Hyb-KAN ViT-B (Mask R-CNN, ViTDet-style, 3x) | 2025-05-07 |
| ResNeSt: Split-Attention Networks | ✓ Link | 50.54 | | | | | | | ResNeSt-200 (Cascade R-CNN, tricks, 3x; official repo) | 2020-04-19 |
| MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models | ✓ Link | 50.5 | | | | | | | tiny-MOAT-0 (IN-1K pretraining, single-scale) | 2022-10-04 |
| Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss | ✓ Link | 50.5 | 67.7 | 55.3 | 34.7 | 53.6 | 64.6 | | Align-DETR (ResNet-50, 12ep) | 2023-04-15 |
| DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection | ✓ Link | 50.5 | 67.5 | 55.1 | 34.3 | 54.6 | 64.5 | 48 | DS-Det (ResNet-50, 12ep) | 2025-07-26 |
| VisionHOPE: Visual Backbones as Self-Modifying Learning Systems | | 50.5 | | | | | | | VisionHOPE-B (Mask R-CNN 1x) | 2026-09-27 |
| VisionHOPE: Visual Backbones as Self-Modifying Learning Systems | | 50.5 | 71.7 | 55.5 | | | | | VisionHOPE-S (Mask R-CNN 3x) | 2026-09-27 |
| FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation | ✓ Link | 50.4 | 68.7 | 54.9 | 32.7 | 53.6 | 65.2 | | H-Deformable-DETR + FeatAug-Flip (ResNet-50, 24ep) | 2023-03-02 |
| Fractional Correspondence Framework in Detection Transformer | | 50.4 | 67.9 | 55.2 | 34.7 | 53.8 | 64.2 | | RTP-DETR (ResNet-50, 12ep) | 2025-03-06 |
| Masked Autoencoders Are Scalable Vision Learners | ✓ Link | 50.3 | | | | | | | MAE (ViT-B, Mask R-CNN) | 2021-11-11 |
| SpectFormer: Frequency and Attention is what you need in a Vision Transformer | | 50.3 | 70.0 | 55.2 | | | | | SpectFormer-H-S (GFL 3x) | 2023-04-13 |
| DAT++: Spatially Dynamic Vision Transformer with Deformable Attention | ✓ Link | 50.2 | 71.5 | 54.0 | 34.7 | 54.6 | 65.3 | 63 | DAT-S++ (RetinaNet 3x) | 2023-09-04 |
| Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens | | 50.2 | 71.4 | 53.9 | 33.2 | 54.5 | 66.3 | 105 | SECViT-L (RetinaNet 1x) | 2024-05-22 |
| PVT v2: Improved Baselines with Pyramid Vision Transformer | ✓ Link | 50.1 | 69.5 | 54.9 | | | | | Sparse R-CNN (PVTv2-B2) | 2021-06-25 |
| GlobalMamba: Global Image Serialization for Vision Mamba | ✓ Link | 50.1 | | 54.9 | | | | 70 | GlobalMamba-S (Mask R-CNN 3x) | 2024-10-14 |
| UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale | ✓ Link | 50.1 | 71.0 | 54.8 | | | | 50.0 | UniConvNet-T (Mask R-CNN 3x) | 2025-08-12 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 50.0 | | | | | | | Pix2seq (ViT-L) | 2021-09-22 |
| V2M: Visual 2-Dimensional Mamba for Image Representation Learning | ✓ Link | 50.0 | 70.9 | 54.8 | | | | 70 | V2M-B (Mask R-CNN 3x) | 2024-10-14 |
| DaViT: Dual Attention Vision Transformers | ✓ Link | 49.9 | | | | | | | DaViT-T (Mask R-CNN, 36 epochs) | 2022-04-07 |
| FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation | ✓ Link | 49.9 | 68.3 | 54.7 | 32.5 | 52.9 | 65.1 | | Deformable-DETR w/ tricks + FeatAug-FC (ResNet-50, 24ep) | 2023-03-02 |
| VMamba: Visual State Space Model | ✓ Link | 49.9 | 70.9 | 54.7 | | | | 70 | VMamba-S (Mask R-CNN 3x) | 2024-01-18 |
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | | 49.9 | | | | | | 114 | OverLoCK-B (Mask R-CNN 1x) | 2025-02-27 |
| Enhanced Training of Query-Based Object Detection via Selective Query Recollection | ✓ Link | 49.8 | 68.8 | 54.0 | 32.0 | 53.4 | 65.1 | | SQR-AdaMixer (ResNet-101) | 2022-12-15 |
| Bottleneck Transformers for Visual Recognition | | 49.7 | 71.3 | 54.6 | | | | | BoTNet 200 (Mask R-CNN, single scale, 72 epochs) | 2021-01-27 |
| Dynamic Head: Unifying Object Detection Heads with Attentions | ✓ Link | 49.7 | 68.0 | 54.3 | 33.3 | 54.2 | 64.2 | | DyHead (Swin-T, 2x) | 2021-06-15 |
| Language-aware Multiple Datasets Detection Pretraining for DETRs | | 49.7 | 67.6 | 53.8 | 32.2 | 52.6 | 64.9 | 50 | METR (ResNet-50, 50ep) | 2023-04-07 |
| Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics | ✓ Link | 49.7 | 71.3 | 54.6 | | | | | VNCT-T (Mask R-CNN 3x) | 2026-07-03 |
| Spectral-Adaptive Modulation Networks for Visual Perception | ✓ Link | 49.6 | 71.3 | 54.8 | | | | | SPANetV2-S18-hybrid (Mask R-CNN 3x) | 2025-03-31 |
| Bottleneck Transformers for Visual Recognition | | 49.5 | 71 | 54.2 | | | | | BoTNet 152 (Mask R-CNN, single scale, 72 epochs) | 2021-01-27 |
| DN-DETR: Accelerate DETR Training by Introducing Query DeNoising | ✓ Link | 49.5 | 67.6 | 53.8 | 31.3 | 52.6 | 65.4 | 47 | DN-Deformable-DETR-R50++ | 2022-03-02 |
| Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN | ✓ Link | 49.4 | | | | | | | A2MIM (ViT-B, Mask R-CNN) | 2022-05-27 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 49.4 | 70.7 | 54.1 | | | | 102 | MogaNet-L (Mask R-CNN 1x) | 2022-11-07 |
| Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation | | 49.4 | 66.9 | 53.9 | 31.8 | 52.3 | 64.6 | 58 | DINO + SACQ (ResNet-50, 12ep) | 2024-05-06 |
| Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders | ✓ Link | 49.4 | 67.9 | 53.9 | 33.1 | 52.6 | 64.0 | | Deformable-DETR++ + Dual-R (ResNet-50, 24ep) | 2025-12-15 |
| VisionHOPE: Visual Backbones as Self-Modifying Learning Systems | | 49.4 | 70.6 | 54.5 | | | | | VisionHOPE-T (Mask R-CNN 3x) | 2026-09-27 |
| GlobalMamba: Global Image Serialization for Vision Mamba | ✓ Link | 49.3 | 71.4 | 54.2 | | | | 108 | GlobalMamba-B (Mask R-CNN 1x) | 2024-10-14 |
| ViDT: An Efficient and Effective Fully Transformer-based Object Detector | ✓ Link | 49.2 | 69.4 | 53.1 | 30.6 | 52.6 | 66.9 | | ViDT (Swin-base, 50ep) | 2021-10-08 |
| DAT++: Spatially Dynamic Vision Transformer with Deformable Attention | ✓ Link | 49.2 | 70.3 | 53.0 | 32.7 | 53.4 | 64.7 | 34 | DAT-T++ (RetinaNet 3x) | 2023-09-04 |
| VMamba: Visual State Space Model | ✓ Link | 49.2 | 71.4 | 54.0 | | | | 108 | VMamba-B (Mask R-CNN 1x) | 2024-01-18 |
| Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement | ✓ Link | 49.2 | 67.1 | 53.8 | 32.7 | 53.0 | 63.1 | 56.1 | Salience DETR (ResNet-50, 12ep) | 2024-03-24 |
| Recurrent Glimpse-based Decoder for Detection with Transformer | ✓ Link | 49.1 | 67.5 | 53.1 | 30 | 52.6 | 65 | | REGO-Deformable DETR-X101 | 2021-12-09 |
| Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion | ✓ Link | 49.1 | 67.2 | 53.2 | 30.5 | 52.6 | 64.7 | 55 | SAM-DETR++ w/ SMCA+DN (ResNet-50, 50ep) | 2022-07-28 |
| HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs | | 49.1 | 71.0 | 54.3 | | | | 69.4 | HIRI-ViT-B (Mask R-CNN 1x, 1600px input) | 2024-03-18 |
| EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm | ✓ Link | 49.0 | 70.3 | 53.6 | | | | 68 | EATFormer-Base (Mask R-CNN 3x) | 2022-06-19 |
| V2M: Visual 2-Dimensional Mamba for Image Representation Learning | ✓ Link | 49.0 | 70.6 | 53.5 | | | | 50 | V2M-S (Mask R-CNN 3x) | 2024-10-14 |
| GlobalMamba: Global Image Serialization for Vision Mamba | ✓ Link | 49.0 | 70.5 | 53.7 | | | | 50 | GlobalMamba-T (Mask R-CNN 3x) | 2024-10-14 |
| GlobalMamba: Global Image Serialization for Vision Mamba | ✓ Link | 49.0 | 70.5 | 53.5 | | | | 70 | GlobalMamba-S (Mask R-CNN 1x) | 2024-10-14 |
| Enhanced Training of Query-Based Object Detection via Selective Query Recollection | ✓ Link | 48.9 | 67.5 | 53.2 | 32.0 | 51.8 | 63.7 | | SQR-AdaMixer (ResNet-50) | 2022-12-15 |
| V2M: Visual 2-Dimensional Mamba for Image Representation Learning | ✓ Link | 48.9 | 70.2 | 53.6 | | | | 70 | V2M-B (Mask R-CNN 1x) | 2024-10-14 |
| VMamba: Visual State Space Model | ✓ Link | 48.8 | 70.4 | 53.5 | | | | 50 | VMamba-T (Mask R-CNN 3x) | 2024-01-18 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 48.7 | 69.5 | 52.6 | 31.5 | 53.4 | 62.7 | 92 | MogaNet-L (RetinaNet 1x) | 2022-11-07 |
| VMamba: Visual State Space Model | ✓ Link | 48.7 | 70.0 | 53.4 | | | | 70 | VMamba-S (Mask R-CNN 1x) | 2024-01-18 |
| Rethinking ImageNet Pre-training | | 48.6 | 66.8 | 52.9 | | | | | Mask R-CNN (ResNeXt-152-FPN, cascade) | 2018-11-21 |
| CenterMask : Real-Time Anchor-Free Instance Segmentation | ✓ Link | 48.6 | | | | | | | CenterMask2 (VoVNetV2-99, 3x, TTA; official repo) | 2019-11-15 |
| USB: Universal-Scale Object Detection Benchmark | ✓ Link | 48.6 | 67.1 | 52.7 | 30.1 | 53.0 | 63.8 | | UniverseNet-20.08d (Res2Net-50, DCN) | 2021-03-25 |
| BiFormer: Vision Transformer with Bi-Level Routing Attention | ✓ Link | 48.6 | 70.5 | 53.8 | | | | | BiFormer-B (Mask R-CNN 1x) | 2023-03-15 |
| XCiT: Cross-Covariance Image Transformers | ✓ Link | 48.5 | | | | | | | XCiT-M24/8 | 2021-06-17 |
| DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention | ✓ Link | 48.5 | 70.2 | 53.3 | | | | | DeBiFormer-B (Mask R-CNN 1x) | 2024-10-11 |
| HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs | | 48.4 | 69.7 | 51.9 | 33.4 | 52.6 | 63.4 | 59.8 | HIRI-ViT-B (RetinaNet 1x, 1600px input) | 2024-03-18 |
| ELSA: Enhanced Local Self-Attention for Vision Transformer | ✓ Link | 48.3 | 70.4 | 52.9 | | | | | ELSA-S (Mask RCNN) | 2021-12-23 |
| Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer | | 48.3 | 69.8 | 52.7 | | | | 51.7 | Hyb-KAN ViT-S (Mask R-CNN, ViTDet-style, 3x) | 2025-05-07 |
| LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking | | 48.2 | 67.5 | 52.6 | 31.7 | 51.3 | 63.8 | 79 | AdaMixer + LORS (ResNet-101, 3x) | 2024-03-07 |
| XCiT: Cross-Covariance Image Transformers | ✓ Link | 48.1 | | | | | | | XCiT-S24/8 | 2021-06-17 |
| GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond | ✓ Link | 47.9 | 66.9 | 52.2 | | | | | GCNet (ResNeXt-101 + DCN + cascade + GC r16) | 2019-04-25 |
| Vision Transformer with Deformable Attention | ✓ Link | 47.9 | 69.6 | 51.2 | 32.3 | 51.8 | 63.4 | | DAT-S (RetinaNet 3x) | 2022-01-03 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 47.9 | 70.0 | 52.7 | | | | 63 | MogaNet-B (Mask R-CNN 1x) | 2022-11-07 |
| BiFormer: Vision Transformer with Bi-Level Routing Attention | ✓ Link | 47.8 | 69.8 | 52.3 | | | | | BiFormer-S (Mask R-CNN 1x) | 2023-03-15 |
| Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics | ✓ Link | 47.8 | 70.3 | 52.6 | | | | | VNCT-T (Mask R-CNN 1x) | 2026-07-03 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 47.7 | 68.9 | 51.0 | 30.5 | 52.2 | 61.7 | 54 | MogaNet-B (RetinaNet 1x) | 2022-11-07 |
| MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection | ✓ Link | 47.6 | | | 30.2 | 51.8 | 60.8 | | MAE-DET-L + GFLV2 | 2021-11-26 |
| Language-aware Multiple Datasets Detection Pretraining for DETRs | | 47.6 | 65.6 | 52.0 | 29.8 | 50.9 | 62.6 | 50 | METR (ResNet-50, 12ep) | 2023-04-07 |
| LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking | | 47.6 | 66.6 | 52.0 | 31.1 | 50.2 | 62.5 | 60 | AdaMixer + LORS (ResNet-50, 3x) | 2024-03-07 |
| Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation | | 47.6 | 65.7 | 51.9 | 30.0 | 50.5 | 62.4 | 58 | DAB-Deformable-DETR + SACQ (ResNet-50, 50ep) | 2024-05-06 |
| ViDT: An Efficient and Effective Fully Transformer-based Object Detector | ✓ Link | 47.5 | 67.7 | 51.4 | 29.2 | 50.7 | 64.8 | 61 | ViDT (Swin-small, 50ep) | 2021-10-08 |
| Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion | ✓ Link | 47.5 | 66.5 | 51.3 | 29.3 | 50.8 | 62.7 | 55 | SAM-DETR++ (ResNet-50, 50ep) | 2022-07-28 |
| DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention | ✓ Link | 47.5 | 69.7 | 52.1 | | | | | DeBiFormer-S (Mask R-CNN 1x) | 2024-10-11 |
| Rethinking ImageNet Pre-training | | 47.4 | | | | | | | Mask R-CNN (ResNet-101-FPN, GN, Cascade) | 2018-11-21 |
| EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm | ✓ Link | 47.4 | 69.3 | 51.9 | | | | 44 | EATFormer-Small (Mask R-CNN 3x) | 2022-06-19 |
| MambaOut: Do We Really Need Mamba for Vision? | ✓ Link | 47.4 | 69.1 | 52.4 | | | | 65 | MambaOut-Small (Mask R-CNN 1x) | 2024-05-13 |
| MambaOut: Do We Really Need Mamba for Vision? | ✓ Link | 47.4 | 69.3 | 52.2 | | | | 100 | MambaOut-Base (Mask R-CNN 1x) | 2024-05-13 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 47.3 | | | | | | | Pix2seq (R50-C4) | 2021-09-22 |
| Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation | | 47.3 | 66.7 | 51.2 | 30.6 | 50.0 | 62.6 | 51 | Two-stage Deformable-DETR + SACQ (ResNet-50, 50ep) | 2024-05-06 |
| EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm | ✓ Link | 47.2 | 69.4 | 52.1 | | | | 68 | EATFormer-Base (Mask R-CNN 1x) | 2022-06-19 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 47.1 | | | | | | | Pix2seq (ViT-B) | 2021-09-22 |
| DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention | ✓ Link | 47.1 | 68.2 | 50.2 | 30.3 | 51.1 | 63.0 | | DeBiFormer-B (RetinaNet 1x) | 2024-10-11 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 47.0 | | | 28.8 | 50.3 | 62.2 | | HTC (HRNetV2p-W48) | 2019-08-20 |
| Augmenting Convolutional networks with attention-based aggregation | ✓ Link | 47.0 | | | | | | | PatchConvNet-S120 (Mask R-CNN) | 2021-12-27 |
| SpectFormer: Frequency and Attention is what you need in a Vision Transformer | | 46.9 | 68.8 | 51.8 | | | | | SpectFormer-B (Mask R-CNN) | 2023-04-13 |
| DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model | | 46.9 | | | 28.8 | 49.5 | 62.9 | | DI-MaskDINO (ResNet-50, 12ep) | 2024-10-22 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 46.8 | | | | | | | RPDet (ResNeXt-101-DCN, multi-scale) | 2019-04-25 |
| Cross Resolution Encoding-Decoding For Detection Transformers | ✓ Link | 46.8 | 66.8 | 50.5 | 27.4 | 50.7 | 64.0 | 45 | DN-DETR + CRED-OO (ResNet-50, 50ep) | 2024-10-05 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 46.7 | 68.0 | 51.3 | | | | 45 | MogaNet-S (Mask R-CNN 1x) | 2022-11-07 |
| DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR | ✓ Link | 46.6 | 67 | 50.2 | 28.1 | 50.5 | 64.1 | 63 | DAB-DETR-DC5-R101 | 2022-01-28 |
| LBMamba: Locally Bi-directional Mamba | ✓ Link | 46.6 | 65.2 | 50.5 | 27.0 | 50.6 | 63.6 | 74 | LBVim-300 (Cascade Mask R-CNN) | 2025-06-19 |
| Rethinking ImageNet Pre-training | | 46.4 | 67.1 | 51.1 | | | | | Mask R-CNN (ResNeXt-152-FPN) | 2018-11-21 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 46.4 | | | | | | | RPDet (ResNet-101-DCN, multi-scale) | 2019-04-25 |
| Sparse R-CNN: End-to-End Object Detection with Learnable Proposals | ✓ Link | 46.4 | 64.6 | 49.5 | 28.3 | 48.3 | 61.6 | | Sparse R-CNN (ResNet-101, learnable proposals, random crop aug, FPN) | 2020-11-25 |
| Augmenting Convolutional networks with attention-based aggregation | ✓ Link | 46.4 | | | | | | | PatchConvNet-S60 (Mask R-CNN) | 2021-12-27 |
| AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs | | 46.4 | 68.5 | 51.3 | | | | 32.3 | AttentionViG-B (Mask R-CNN 1x) | 2025-09-29 |
| PVT v2: Improved Baselines with Pyramid Vision Transformer | ✓ Link | 46.3 | 64.3 | 50.5 | | | | | Cascade Mask R-CNN (ResNet-50, 3x; reported by PVTv2) | 2021-06-25 |
| Cross Resolution Encoding-Decoding For Detection Transformers | ✓ Link | 46.2 | 65.8 | 49.8 | 26.8 | 50.0 | 63.5 | 45 | DN-DETR + CRED (ResNet-50, 50ep) | 2024-10-05 |
| HoughNet: Integrating near and long-range evidence for bottom-up object detection | ✓ Link | 46.1 | 64.6 | 50.3 | 30.0 | 48.8 | 59.7 | | HoughNet (HG-104, MS) | 2020-07-05 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 46.0 | | | 27.5 | 48.9 | 60.1 | | Cascade Mask R-CNN (HRNetV2p-W48) | 2019-08-20 |
| SpectFormer: Frequency and Attention is what you need in a Vision Transformer | | 46.0 | 66.4 | 49.7 | 29.5 | 49.7 | 61.1 | | SpectFormer-B (RetinaNet) | 2023-04-13 |
| Bottleneck Transformers for Visual Recognition | | 45.9 | | | | | | | BoTNet 50 (72 epochs) | 2021-01-27 |
| Conditional DETR for Fast Training Convergence | ✓ Link | 45.9 | 66.8 | 49.5 | 27.2 | 50.3 | 63.3 | 63 | Conditional DETR-DC5-R101 | 2021-08-13 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 45.8 | 66.6 | 49.0 | 29.1 | 50.1 | 59.8 | 35 | MogaNet-S (RetinaNet 1x) | 2022-11-07 |
| CenterMask : Real-Time Anchor-Free Instance Segmentation | ✓ Link | 45.6 | | | 29.2 | 49.3 | 58.8 | | CenterMask (VoVNetV2-99, 3x; official repo) | 2019-11-15 |
| DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention | ✓ Link | 45.6 | 66.6 | 48.9 | 28.7 | 49.3 | 61.6 | | DeBiFormer-S (RetinaNet 1x) | 2024-10-11 |
| Cross Resolution Encoding-Decoding For Detection Transformers | ✓ Link | 45.4 | 64.9 | 49.4 | 27.0 | 48.5 | 62.2 | 45 | DAB-DETR + CRED (ResNet-50, 50ep) | 2024-10-05 |
| LBMamba: Locally Bi-directional Mamba | ✓ Link | 45.4 | 63.8 | 49.3 | 25.5 | 49.4 | 62.4 | 66 | LBVim-Ti (Cascade Mask R-CNN) | 2025-06-19 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 45.3 | | | 27.0 | 48.4 | 59.5 | | HTC (HRNetV2p-W32) | 2019-08-20 |
| Conditional DETR for Fast Training Convergence | ✓ Link | 45.1 | 65.4 | 48.5 | 25.3 | 49 | 62.2 | 44 | Conditional DETR-DC5-R50 | 2021-08-13 |
| Anchor DETR: Query Design for Transformer-Based Object Detection | ✓ Link | 45.1 | 65.7 | 48.8 | 25.8 | 49.4 | 61.6 | | Anchor DETR-DC5-R101 | 2021-09-15 |
| MambaOut: Do We Really Need Mamba for Vision? | ✓ Link | 45.1 | 67.3 | 49.6 | | | | 43 | MambaOut-Tiny (Mask R-CNN 1x) | 2024-05-13 |
| Non-local Neural Networks | ✓ Link | 45.0 | 67.8 | 48.9 | | | | | Mask R-CNN (ResNeXt-152 + 1 NL) | 2017-11-21 |
| Sparse R-CNN: End-to-End Object Detection with Learnable Proposals | ✓ Link | 45.0 | 63.4 | 48.2 | 26.9 | 47.2 | 59.5 | | Sparse R-CNN (ResNet-50, learnable proposals, random crop aug, FPN) | 2020-11-25 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 45.0 | 63.2 | 48.6 | 28.2 | 48.9 | 60.4 | | Pix2seq (R101-DC5) | 2021-09-22 |
| Attentive Normalization | ✓ Link | 44.9 | 66.2 | 49.1 | | | | | Mask R-CNN-FPN (AOGNet-40M) | 2019-08-04 |
| CenterMask : Real-Time Anchor-Free Instance Segmentation | ✓ Link | 44.9 | | | | | | | Mask R-CNN (VoVNetV2-99, 3x; CenterMask2 repo) | 2019-11-15 |
| End-to-End Object Detection with Transformers | ✓ Link | 44.9 | 64.7 | 47.7 | 23.7 | 49.5 | 62.3 | | DETR-DC5 (ResNet-101) | 2020-05-26 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 44.8 | | | | | | | RPDet (ResNet-101-DCN, multi-scale train) | 2019-04-25 |
| Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing | ✓ Link | 44.8 | 64.3 | 48.9 | 26.6 | 48.3 | 59.6 | | R3-CNN (ResNet-50-FPN, DCN) | 2021-04-03 |
| ViDT: An Efficient and Effective Fully Transformer-based Object Detector | ✓ Link | 44.8 | 64.5 | 48.7 | 25.9 | 47.6 | 62.1 | 38 | ViDT (Swin-tiny, 50ep) | 2021-10-08 |
| Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion | ✓ Link | 44.8 | 62.6 | 47.9 | 26.7 | 48.2 | 60.9 | 55 | SAM-DETR++ w/ SMCA+DN (ResNet-50, 12ep) | 2022-07-28 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 44.6 | 62.7 | 48.7 | 26.3 | 48.1 | 58.5 | | Cascade R-CNN (HRNetV2p-W48) | 2019-08-20 |
| CenterMask : Real-Time Anchor-Free Instance Segmentation | ✓ Link | 44.6 | | | 27.7 | 48.3 | 57.3 | | CenterMask (VoVNetV2-57, 3x; official repo) | 2019-11-15 |
| RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations | ✓ Link | 44.6 | 66.3 | 49.0 | | | | | RecNeXt-M5 (Mask R-CNN) | 2024-12-27 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 44.5 | | | | | | | RPDet (ResNeXt-101-DCN) | 2019-04-25 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 44.5 | | | 26.1 | 47.9 | 58.5 | | Cascade Mask R-CNN (HRNetV2p-W32) | 2019-08-20 |
| PVT v2: Improved Baselines with Pyramid Vision Transformer | ✓ Link | 44.5 | 63.0 | 48.3 | | | | | GFL (ResNet-50, 3x; reported by PVTv2) | 2021-06-25 |
| Conditional DETR for Fast Training Convergence | ✓ Link | 44.5 | 65.6 | 47.5 | 23.6 | 48.4 | 63.6 | 63 | Conditional DETR-R101 | 2021-08-13 |
| CenterMask : Real-Time Anchor-Free Instance Segmentation | ✓ Link | 44.4 | | | 26.7 | 47.7 | 57.1 | | CenterMask (X-101-32x8d, 3x; official repo) | 2019-11-15 |
| Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing | ✓ Link | 44.3 | 64.1 | 48.4 | 27 | 47.1 | 58.9 | | R3-CNN (ResNet-50-FPN, GC-Net) | 2021-04-03 |
| Anchor DETR: Query Design for Transformer-Based Object Detection | ✓ Link | 44.2 | 64.7 | 47.5 | 24.7 | 48.2 | 60.6 | | Anchor DETR-DC5-R50 | 2021-09-15 |
| SCSA: Exploring the Synergistic Effects Between Spatial and Channel Attention | | 44.2 | 63.1 | 48.2 | 26.0 | 48.2 | 57.5 | 88.52 | Cascade R-CNN + SCSA (ResNet-101, 1x) | 2024-07-06 |
| Sparse R-CNN: End-to-End Object Detection with Learnable Proposals | ✓ Link | 44.1 | 62.1 | 47.2 | 26.1 | 46.3 | 59.7 | | Sparse R-CNN (ResNet-101, FPN) | 2020-11-25 |
| DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR | ✓ Link | 44.1 | 64.7 | 47.2 | 24.1 | 48.2 | 62.9 | 63 | DAB-DETR-R101 | 2022-01-28 |
| Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond | | 44.07 | 64.48 | 48.40 | 26.73 | 48.11 | 56.80 | | Faster R-CNN + LossMix (ResNet-101-FPN, 3x) | 2023-03-18 |
| End-to-End Object Detection with Transformers | ✓ Link | 44 | 63.9 | 47.8 | 27.2 | 48.1 | 56 | | Faster RCNN-R101-FPN+ | 2020-05-26 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 43.7 | 61.7 | 47.7 | 25.6 | 46.5 | 57.4 | | Cascade R-CNN (HRNetV2p-W32) | 2019-08-20 |
| Micro-Batch Training with Batch-Channel Normalization and Weight Standardization | ✓ Link | 43.6 | 64.4 | 47.9 | 25.6 | 47.5 | 57.4 | | Mask R-CNN-FPN (ResNet-101, GN+WS+BCN) | 2019-03-25 |
| PVT v2: Improved Baselines with Pyramid Vision Transformer | ✓ Link | 43.5 | 61.9 | 47.0 | | | | | ATSS (ResNet-50, 3x; reported by PVTv2) | 2021-06-25 |
| RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations | ✓ Link | 43.5 | 64.9 | 47.7 | | | | | RecNeXt-M4 (Mask R-CNN) | 2024-12-27 |
| AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs | | 43.5 | 65.8 | 47.6 | | | | 12.3 | AttentionViG-S (Mask R-CNN 1x) | 2025-09-29 |
| Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions | ✓ Link | 43.4 | 63.6 | 46.1 | 26.1 | 46.0 | 59.5 | | PVT-Large (RetinaNet 3x,MS) | 2021-02-24 |
| Bottom-up Object Detection by Grouping Extreme and Center Points | ✓ Link | 43.3 | 59.6 | 46.8 | 25.7 | 46.6 | 59.4 | | ExtremeNet (Hourglass-104, multi-scale) | 2019-01-23 |
| Hybrid Task Cascade for Instance Segmentation | ✓ Link | 43.2 | 59.4 | 40.7 | 20.3 | 40.9 | 52.3 | | HTC (cascade) | 2019-01-22 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 43.2 | 61.0 | 46.1 | 26.6 | 47 | 58.6 | | Pix2seq (R50-DC5 ) | 2021-09-22 |
| Deformable ConvNets v2: More Deformable, Better Results | | 43.1 | | | | | | | Mask R-CNN (ResNet-101, DCNv2) | 2018-11-27 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 43.1 | | | 26.6 | 46.0 | 56.9 | | HTC (HRNetV2p-W18) | 2019-08-20 |
| HoughNet: Integrating near and long-range evidence for bottom-up object detection | ✓ Link | 43.0 | 62.2 | 46.9 | 25.5 | 47.6 | 55.8 | | HoughNet (HG-104) | 2020-07-05 |
| Conditional DETR for Fast Training Convergence | ✓ Link | 43 | 64 | 45.7 | 22.7 | 46.7 | 61.5 | 44 | Conditional DETR-R50 | 2021-08-13 |
| Scaling Graph Convolutions for Mobile Vision | ✓ Link | 43.0 | 64.9 | 47.1 | | | | 27.7 | MobileViGv2-B (Mask R-CNN 1x) | 2024-06-09 |
| Sparse R-CNN: End-to-End Object Detection with Learnable Proposals | ✓ Link | 42.8 | 61.2 | 45.7 | 26.7 | 44.6 | 57.6 | | Sparse R-CNN (ResNet-50, FPN) | 2020-11-25 |
| X-volution: On the unification of convolution and self-attention | | 42.8 | 64 | 46.4 | 26.9 | 46 | 55 | | Faster R-CNN (FPN, X-volution) | 2021-06-04 |
| UniHead: Unifying Multi-Perception for Detection Heads | ✓ Link | 42.8 | 61.0 | 46.2 | 24.8 | 46.8 | 56.8 | | UniHead (ResNet-50, 1x) | 2023-09-23 |
| Cascade R-CNN: Delving into High Quality Object Detection | ✓ Link | 42.7 | 61.6 | 46.6 | 23.8 | 46.2 | 57.4 | | Cascade R-CNN (ResNet-101-FPN+, cascade) | 2017-12-03 |
| Dynamic Feature Pyramid Networks for Object Detection | | 42.7 | | | | | | | Cascade R-CNN + DyFPN (ResNet-101, 1x) | 2020-12-01 |
| CornerNet-Lite: Efficient Keypoint Based Object Detection | ✓ Link | 42.6 | | | 25.5 | 44.3 | 58.4 | | CornerNet-Saccade (Hourglass-54) | 2019-04-18 |
| Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions | ✓ Link | 42.6 | 63.7 | 45.4 | 25.8 | 46.0 | 58.4 | | PVT-Large (RetinaNet 1x) | 2021-02-24 |
| Pix2seq: A Language Modeling Framework for Object Detection | ✓ Link | 42.6 | | | | | | | Pix2seq (R50) | 2021-09-22 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 42.6 | 64.0 | 46.4 | | | | 25 | MogaNet-T (Mask R-CNN 1x) | 2022-11-07 |
| S2AFormer: Strip Self-Attention for Efficient Vision Transformer | | 42.6 | 64.5 | 46.9 | | | | 42 | S2AFormer-M (Mask R-CNN 1x) | 2025-05-28 |
| Scaling Graph Convolutions for Mobile Vision | ✓ Link | 42.5 | 63.9 | 46.3 | | | | 15.4 | MobileViGv2-M (Mask R-CNN 1x) | 2024-06-09 |
| Group Normalization | ✓ Link | 42.3 | 62.8 | 46.2 | | | | | Mask R-CNN (ResNet-101-FPN, GroupNorm, long) | 2018-03-22 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 42.3 | | | 25.0 | 45.4 | 54.9 | | Mask R-CNN (HRNetV2p-W32) | 2019-08-20 |
| Rethinking and Improving Relative Position Encoding for Vision Transformer | ✓ Link | 42.3 | | | | | | | DETR-ResNet50 with iRPE-K (300 epochs) | 2021-07-29 |
| UniHead: Unifying Multi-Perception for Detection Heads | ✓ Link | 42.3 | 61.1 | 45.6 | 24.5 | 46.0 | 55.3 | 31.81 | GFL + UniHead (ResNet-50, 1x) | 2023-09-23 |
| Scale-Aware Trident Networks for Object Detection | ✓ Link | 42 | 63.5 | 45.5 | 24.9 | 47 | 56.9 | | TridentNet (ResNet-101) | 2019-01-07 |
| Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing | ✓ Link | 42 | 61 | 46.3 | 24.5 | 45.2 | 55.7 | | R3-CNN (ResNet-50-FPN) | 2021-04-03 |
| Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond | | 41.82 | 62.51 | 45.81 | 25.04 | 45.48 | 54.03 | | Faster R-CNN + LossMix (ResNet-50-FPN, 3x) | 2023-03-18 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 41.8 | 62.8 | 45.9 | | 44.7 | 54.6 | | Faster R-CNN (HRNetV2p-W48) | 2019-08-20 |
| Deformable ConvNets v2: More Deformable, Better Results | | 41.7 | | | 22.2 | 45.8 | 58.7 | | Faster R-CNN (ResNet-101, DCNv2) | 2018-11-27 |
| LIP: Local Importance-based Pooling | ✓ Link | 41.7 | 63.6 | 45.6 | 25.2 | 45.8 | | | Faster R-CNN (LIP-ResNet-101) | 2019-08-12 |
| RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations | ✓ Link | 41.7 | 63.4 | 45.4 | | | | | RecNeXt-M3 (Mask R-CNN) | 2024-12-27 |
| S2AFormer: Strip Self-Attention for Efficient Vision Transformer | | 41.7 | 62.4 | 44.5 | 25.8 | 44.6 | 55.4 | 32 | S2AFormer-M (RetinaNet 1x) | 2025-05-28 |
| Feature Selective Anchor-Free Module for Single-Shot Object Detection | | 41.6 | 62.4 | | | | | | FSAF (ResNeXt-101, anchor-based branches) | 2019-03-02 |
| SCSA: Exploring the Synergistic Effects Between Spatial and Channel Attention | | 41.5 | 62.9 | 45.4 | 24.6 | 45.3 | 53.7 | 60.88 | Faster R-CNN + SCSA (ResNet-101, 1x) | 2024-07-06 |
| CornerNet-Lite: Efficient Keypoint Based Object Detection | ✓ Link | 41.4 | | | 23.8 | 43.5 | 57.1 | | CornerNet-Saccade (Hourglass-104) | 2019-04-18 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 41.4 | 61.5 | 44.4 | 25.1 | 45.7 | 53.6 | 14 | MogaNet-T (RetinaNet 1x) | 2022-11-07 |
| Grid R-CNN | ✓ Link | 41.3 | 60.3 | 44.4 | 23.4 | 45.8 | 54.1 | | Grid R-CNN (ResNet-101-FPN) | 2018-11-29 |
| CenterNet: Keypoint Triplets for Object Detection | ✓ Link | 41.3 | 59.2 | 43.9 | 23.6 | 43.8 | 55.8 | | CenterNet511 (Hourglass-52) | 2019-04-17 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 41.3 | 59.2 | 44.9 | 23.7 | 44.2 | 54.1 | | Cascade R-CNN (HRNetV2p-W18) | 2019-08-20 |
| RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free | ✓ Link | 41.1 | 60.2 | 44.1 | | | | | RetinaMask (ResNet-101-FPN) | 2019-01-10 |
| MetaFormer Is Actually What You Need for Vision | ✓ Link | 41.0 | 63.1 | 44.8 | | | | | PoolFormer-S36 (Mask R-CNN) | 2021-11-22 |
| LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection | ✓ Link | 41.0 | 57.9 | 44.3 | 21.9 | 46.1 | 56.8 | 2.4 | LeYOLO-Large (768) | 2024-06-20 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 40.9 | 61.8 | 44.8 | 24.4 | 43.7 | 53.3 | | Faster R-CNN (HRNetV2p-W32) | 2019-08-20 |
| VirTex: Learning Visual Representations from Textual Annotations | ✓ Link | 40.9 | | | | | | | VirTex Mask R-CNN (ResNet-50-FPN) | 2020-06-11 |
| Dynamic Feature Pyramid Networks for Object Detection | | 40.9 | | | | | | | Mask R-CNN + DyFPN (ResNet-101, 1x) | 2020-12-01 |
| PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices | ✓ Link | 40.9 | 57.6 | | | | | 3.30 | PP-PicoDet-L (640) | 2021-11-01 |
| Non-local Neural Networks | ✓ Link | 40.8 | 63.1 | 44.5 | | | | | Mask R-CNN (ResNet-101 + 1 NL) | 2017-11-21 |
| Group Normalization | ✓ Link | 40.8 | 61.6 | 44.4 | | | | | Mask R-CNN (ResNet-50-FPN, GroupNorm, long) | 2018-03-22 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 40.8 | | | | | | | RPDet (ResNet-50, multi-scale train) | 2019-04-25 |
| Rethinking and Improving Relative Position Encoding for Vision Transformer | ✓ Link | 40.8 | | | | | | | DETR-ResNet50 with iRPE-K (150 epochs) | 2021-07-29 |
| LSNet: See Large, Focus Small | ✓ Link | 40.8 | 63.4 | 44.0 | | | | | LSNet-B (Mask R-CNN 1x) | 2025-03-29 |
| A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection | ✓ Link | 40.7 | 60.7 | 43.3 | | | | | Faster R-CNN+aLRP Loss (ResNet-50, 500 scale) | 2020-09-28 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 40.7 | 62.3 | 44.4 | | | | 23 | MogaNet-XT (Mask R-CNN 1x) | 2022-11-07 |
| ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding | | 40.7 | 63.3 | 44.2 | | | | 42.2 | ParFormer-M (Mask R-CNN 1x) | 2024-03-22 |
| Acquisition of Localization Confidence for Accurate Object Detection | ✓ Link | 40.6 | 59.0 | | | | | | IoU-Net (ResNet-101-FPN, IoU-NMS + refinement) | 2018-07-30 |
| Reducing Label Noise in Anchor-Free Object Detection | ✓ Link | 40.5 | 59.5 | 44.2 | 25.4 | 44.7 | 52.3 | | PPDet (ResNet-101-FPN) | 2020-08-03 |
| ViDT: An Efficient and Effective Fully Transformer-based Object Detector | ✓ Link | 40.4 | 59.6 | 43.3 | 23.2 | 42.5 | 55.8 | 16 | ViDT (Swin-nano, 50ep) | 2021-10-08 |
| A lightweight mechanism for vision-transformer-based object detection | | 40.4 | 58.9 | 43.9 | 24.0 | 43.9 | 52.0 | | XFCOS (ResNet-50, 3x) | 2025-05-22 |
| Cascade R-CNN: Delving into High Quality Object Detection | ✓ Link | 40.3 | 59.4 | 43.7 | 22.9 | 43.7 | 54.1 | | Cascade R-CNN (ResNet-50-FPN+) | 2017-12-03 |
| Group Normalization | ✓ Link | 40.3 | 61 | 44 | | | | | Mask R-CNN (ResNet-50-FPN, GroupNorm) | 2018-03-22 |
| Bottom-up Object Detection by Grouping Extreme and Center Points | ✓ Link | 40.3 | 55.1 | 43.7 | 21.6 | 44.0 | 56.1 | | ExtremeNet (Hourglass-104, single-scale) | 2019-01-23 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 40.3 | | | | | | | RPDet (ResNet-101) | 2019-04-25 |
| A novel Region of Interest Extraction Layer for Instance Segmentation | ✓ Link | 40.3 | 62.4 | 44 | 24.2 | 44.4 | 52.5 | | GC-Net + GRoIE (ResNet-50-FPN) | 2020-04-28 |
| A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection | ✓ Link | 40.2 | 60.3 | 42.3 | | | | | RetinaNet+aLRP Loss (ResNet-50, 500 scale) | 2020-09-28 |
| Cross-Iteration Batch Normalization | ✓ Link | 40.1 | 60.5 | 44.1 | | | | | Mask R-CNN (ResNet-101-FPN, CBN) | 2020-02-13 |
| Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN | ✓ Link | 39.8 | | | | | | | A2MIM (ResNet-50-C4, Mask R-CNN 2x) | 2022-05-27 |
| A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection | ✓ Link | 39.7 | 58.8 | 41.5 | | | | | FoveaBox+aLRP Loss (ResNet-50, 500 scale) | 2020-09-28 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 39.7 | 60.0 | 42.4 | 23.8 | 43.6 | 51.7 | 12 | MogaNet-XT (RetinaNet 1x) | 2022-11-07 |
| Grid R-CNN | ✓ Link | 39.6 | 58.3 | 42.4 | 22.6 | 43.8 | 51.5 | | Grid R-CNN (ResNet-50-FPN) | 2018-11-29 |
| Adaptively Connected Neural Networks | ✓ Link | 39.5 | | | | | | | Mask R-CNN (ResNet-50, ACNet) | 2019-04-07 |
| Feature Selective Anchor-Free Module for Single-Shot Object Detection | | 39.3 | 59.2 | | | | | | FSAF (ResNet-101, anchor-based branches) | 2019-03-02 |
| LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection | ✓ Link | 39.3 | 55.7 | 42.5 | 18.8 | 44.1 | 56.1 | | LeYOLO-Medium (640) | 2024-06-20 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 39.2 | | | 23.7 | 41.7 | 51.0 | | Mask R-CNN (HRNetV2p-W18) | 2019-08-20 |
| LSNet: See Large, Focus Small | ✓ Link | 39.2 | 60.0 | 41.5 | 22.1 | 43.0 | 52.9 | | LSNet-B (RetinaNet 1x) | 2025-03-29 |
| Non-local Neural Networks | ✓ Link | 39.0 | 61.1 | 41.9 | | | | | Mask R-CNN (ResNet-50 + 1 NL) | 2017-11-21 |
| FoveaBox: Beyond Anchor-based Object Detector | ✓ Link | 38.9 | 58.4 | 41.5 | 22.3 | 43.5 | 51.7 | | FoveaBox (ResNet-101-FPN, 800x800) | 2019-04-08 |
| FCOS: Fully Convolutional One-Stage Object Detection | ✓ Link | 38.6 | 57.4 | 41.4 | 22.3 | 42.5 | 49.8 | | FCOS (ResNet-50-FPN + improvements) | 2019-04-02 |
| RepPoints: Point Set Representation for Object Detection | ✓ Link | 38.6 | | | | | | | RPDet (ResNet-50) | 2019-04-25 |
| Libra R-CNN: Towards Balanced Learning for Object Detection | ✓ Link | 38.5 | 59.3 | 42.0 | 22.9 | 42.1 | 50.5 | | Libra R-CNN (ResNet-50 FPN) | 2019-04-04 |
| CornerNet: Detecting Objects as Paired Keypoints | ✓ Link | 38.4 | 53.8 | 40.9 | 18.6 | 40.5 | 51.8 | | CornerNet511 (Hourglass-104) | 2018-08-03 |
| A novel Region of Interest Extraction Layer for Instance Segmentation | ✓ Link | 38.4 | 59.9 | 41.7 | 22.9 | 42.1 | 49.7 | | Mask R-CNN (ResNet-50-FPN, GRoIE) | 2020-04-28 |
| LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection | ✓ Link | 38.2 | 54.1 | 41.3 | 17.6 | 42.2 | 55.1 | 1.9 | LeYOLO-Small (640) | 2024-06-20 |
| FoveaBox: Beyond Anchor-based Object Detector | ✓ Link | 38 | 57.8 | 40.2 | 19.5 | 42.2 | 52.7 | | FoveaBox (ResNet-101-FPN, 600x600) | 2019-04-08 |
| Deep High-Resolution Representation Learning for Visual Recognition | ✓ Link | 38.0 | 58.9 | 41.5 | 22.6 | 40.8 | 49.6 | | Faster R-CNN (HRNetV2p-W18) | 2019-08-20 |
| Feature Selective Anchor-Free Module for Single-Shot Object Detection | | 37.9 | 58.0 | | | | | | FSAF (ResNet-101) | 2019-03-02 |
| ELA: Efficient Local Attention for Deep Convolutional Neural Networks | | 37.81 | 57.9 | 40.2 | 18.8 | 42.6 | 52.4 | | YOLOF + ELA-L (1x) | 2024-03-02 |
| A novel Region of Interest Extraction Layer for Instance Segmentation | ✓ Link | 37.5 | 59.2 | 40.6 | 22.3 | 41.5 | 47.8 | | Faster R-CNN (ResNet-50-FPN, GRoIE) | 2020-04-28 |
| Boosting Convolutional Neural Networks with Middle Spectrum Grouped Convolution | ✓ Link | 37.4 | 59.1 | 40.2 | | | | | Faster R-CNN (ResNet-50-MSGC, 1x) | 2023-04-13 |
| ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding | | 37.3 | 59.0 | 40.4 | | | | 25.7 | ParFormer-T (Mask R-CNN 1x) | 2024-03-22 |
| PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices | ✓ Link | 37.2 | 56.0 | | | | | 7.2 | YOLOv5s (640, reported by PP-PicoDet) | 2021-11-01 |
| torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation | ✓ Link | 36.9 | | | | | | | Mask R-CNN (Bottleneck-injected ResNet-50, FPN) | 2020-11-25 |
| Mask R-CNN | ✓ Link | 36.7 | 59.5 | 38.9 | | | | | Mask R-CNN (ResNeXt-101-FPN) | 2017-03-20 |
| FoveaBox: Beyond Anchor-based Object Detector | ✓ Link | 36.0 | 55.2 | 37.9 | 18.6 | 39.4 | 50.5 | | FoveaBox (ResNet-50-FPN, 600x600) | 2019-04-08 |
| Feature Selective Anchor-Free Module for Single-Shot Object Detection | | 35.9 | 55.0 | 37.9 | 19.8 | 39.6 | 48.2 | | FSAF (ResNet-50) | 2019-03-02 |
| torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation | ✓ Link | 35.9 | | | | | | | Faster R-CNN (Bottleneck-injected ResNet-50 and FPN) | 2020-11-25 |
| Gradient Harmonized Single-stage Detector | ✓ Link | 35.8 | 55.5 | 38.1 | 19.6 | 39.6 | 46.7 | | GHM-C + GHM-R (RetinaNet-FPN-ResNet-50, M=30) | 2018-11-13 |
| Generating Positive Bounding Boxes for Balanced Training of Object Detectors | ✓ Link | 35.6 | 55.3 | | | | | | Online Fg Bal. Sampling+Hard Negative Mining (ResNet-50) | 2019-09-21 |
| PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices | ✓ Link | 34.3 | 49.8 | | | | | 2.15 | PP-PicoDet-M (416) | 2021-11-01 |
| M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network | ✓ Link | 34.1 | 53.7 | | 15.9 | 39.5 | 49.3 | | M2Det (ResNet-101, 320x320) | 2018-11-12 |
| Res2Net: A New Multi-scale Backbone Architecture | ✓ Link | 33.7 | 53.6 | | 14 | 38.3 | 51.1 | | Faster R-CNN (Res2Net-50) | 2019-04-02 |
| M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network | ✓ Link | 33.2 | 52.2 | | 15 | 38.2 | 49.1 | | M2Det (VGG-16, 320x320) | 2018-11-12 |
| YOLOX: Exceeding YOLO Series in 2021 | ✓ Link | 32.8 | | | | | | 5.06 | YOLOX-Tiny (416) | 2021-07-18 |
| YOLOX: Exceeding YOLO Series in 2021 | ✓ Link | 25.3 | | | | | | 0.91 | YOLOX-Nano (416) | 2021-07-18 |
| PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices | ✓ Link | 23.5 | | | | | | 0.95 | NanoDet-M (416, reported by PP-PicoDet) | 2021-11-01 |
| You Only Learn One Representation: Unified Network for Multiple Tasks | ✓ Link | | 73.5 | 60.6 | 40.4 | 60.1 | 68.7 | | YOLOR-D6 (1280, single-scale, 31 fps) | 2021-05-10 |
| You Only Learn One Representation: Unified Network for Multiple Tasks | ✓ Link | | 70.6 | 57.4 | 37.4 | 57.3 | 65.2 | | YOLOR-P6 (1280, single-scale, 72 fps) | 2021-05-10 |
| Focal Modulation Networks | ✓ Link | | 70.1 | 55.8 | | | | | FocalNet-T (SRF, Cascade Mask R-CNN) | 2022-03-22 |
| Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing | ✓ Link | | 61.2 | 45.6 | 24.4 | | | | R3-CNN (ResNet-50-FPN, GRoIE) | 2021-04-03 |
| When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism | ✓ Link | | | | | 42.3 | | | Shift-T | 2022-01-26 |