OpenCodePapers

object-detection-on-coco-val2017

Object Detection
Dataset Link
Results over time
Click legend items to toggle metrics. Hover points for model names.
Leaderboard
PaperCodebox APAP50AP75APSAPMAPLParams (M)ModelNameReleaseDate
DINOv3✓ Link66.17100DINOv3 7B/16 (Plain-DETR, frozen backbone, TTA)2025-08-13
Perception Encoder: The best visual embeddings are not at the output of the network✓ Link66.01900PE_spatial (DETA)2025-04-17
DETRs with Collaborative Hybrid Assignments Training✓ Link65.9314Co-DETR2022-11-22
DINOv3✓ Link65.67100DINOv3 7B/16 (Plain-DETR, frozen backbone, no TTA)2025-08-13
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions✓ Link65.0InternImage-H2022-11-10
DETRs with Collaborative Hybrid Assignments Training✓ Link64.7218Co-DETR (Swin-L)2022-11-22
A Strong and Reproducible Object Detector with Only Public Datasets✓ Link64.681.571.450.468.578.5689Focal-Stable-DINO (Focal-Huge, no TTA)2023-04-25
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale✓ Link64.5EVA (Cascade Mask R-CNN, TTA)2022-11-14
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions✓ Link64.3363ViT-CoMer2024-03-12
Focal Modulation Networks✓ Link64.2FocalNet-H (DINO)2022-03-22
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions✓ Link64.2InternImage-XL2022-11-10
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection64.1CP-DETR-L Swin-L(Fine tuning separately in COCO)2024-12-13
Reversible Column Networks✓ Link63.8RevCol-H(DINO)2022-12-22
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection✓ Link63.2DINO (Swin-L)2022-03-07
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection✓ Link63.0Grounding DINO L (1.5x image size)2023-03-09
Swin Transformer V2: Scaling Up Capacity and Resolution✓ Link62.5SwinV2-G (HTC++)2021-11-18
Florence: A New Foundation Model for Computer Vision62Florence-CoSwin-H2021-11-22
General Object Foundation Model for Images and Videos at Scale✓ Link62.0GLEE-Pro2023-12-14
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration✓ Link61.8Fast-iTPN-B (DINO, CLIP-distilled pre-training + Objects365 detection pre-training)2022-11-23
Mr. DETR: Instructive Multi-Route Training for Detection Transformers✓ Link61.879.067.647.765.675.7Mr. DETR-Align (Swin-L, Objects365 pre-training, 5-scale; official repo)2024-12-13
Exploring Plain Vision Transformer Backbones for Object Detection✓ Link61.3ViTDet, ViT-H Cascade (multiscale)2022-03-30
Grounded Language-Image Pre-training✓ Link60.8GLIP (Swin-L, multi-scale)2021-12-07
End-to-End Semi-Supervised Object Detection with Soft Teacher✓ Link60.7Soft Teacher + Swin-L (HTC++, multi-scale)2021-06-16
Universal Instance Perception as Object Discovery and Retrieval✓ Link60.677.566.745.164.875.3UNINEXT-H2023-03-12
Vision Transformer Adapter for Dense Predictions✓ Link60.5ViT-Adapter-L (HTC++, BEiTv2 pretrain, multi-scale)2022-05-17
Exploring Plain Vision Transformer Backbones for Object Detection✓ Link60.4ViTDet, ViT-H Cascade2022-03-30
General Object Foundation Model for Images and Videos at Scale✓ Link60.4GLEE-Plus2023-12-14
Dynamic Head: Unifying Object Detection Heads with Attentions✓ Link60.3DyHead (Swin-L, multi scale, self-training)2021-06-15
Vision Transformer Adapter for Dense Predictions✓ Link60.2ViT-Adapter-L (HTC++, BEiT pretrain, multi-scale)2022-05-17
End-to-End Semi-Supervised Object Detection with Soft Teacher✓ Link60.1Soft Teacher+Swin-L(HTC++, single scale)2021-06-16
Parameter-Inverted Image Pyramid Networks✓ Link60.079.065.4PIIP-H6B (DINO)2024-06-06
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation✓ Link59.842.263.874.9SimPLR (ViT-H, MAE IN-1K pre-training, 1024px)2023-10-09
CBNet: A Composite Backbone Network Architecture for Object Detection✓ Link59.6CBNetV2 (Dual-Swin-L HTC, multi-scale)2021-07-01
Could Giant Pretrained Image Models Extract Universal Representations?59.3Frozen Backbone, SwinV2-G-ext22K (HTC)2022-11-03
HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions✓ Link59.2HorNet-L2022-07-28
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link59.2MOAT-3 (IN-22K pretraining, single-scale)2022-10-04
CBNet: A Composite Backbone Network Architecture for Object Detection✓ Link59.1CBNetV2 (Dual-Swin-L HTC, multi-scale)2021-07-01
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration✓ Link58.8Fast-iTPN-L (DINO 12ep, CLIP-distilled pre-training)2022-11-23
ScaleDet: A Scalable Multi-Dataset Object Detector58.8ScaleDet-B (Swin-B, multi-dataset training)2023-06-08
Focal Self-attention for Local-Global Interactions in Vision Transformers✓ Link58.7Focal-L (DyHead, multi-scale)2021-07-01
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection✓ Link58.7MViTv2-L (Cascade Mask R-CNN, multi-scale, IN21k pre-train)2021-12-02
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation✓ Link58.740.463.274.8SimPLR (ViT-L, BEiTv2 IN-21K pre-training, 1024px)2023-10-09
Mr. DETR: Instructive Multi-Route Training for Detection Transformers✓ Link58.776.564.042.262.975.2Mr. DETR++ (Swin-L, 12ep)2024-12-13
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link58.5MOAT-2 (IN-22K pretraining, single-scale)2022-10-04
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation✓ Link58.542.262.573.4SimPLR (ViT-L, MAE IN-1K pre-training, 1024px)2023-10-09
Dynamic Head: Unifying Object Detection Heads with Attentions✓ Link58.4DyHead (Swin-L, multi scale)2021-06-15
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration✓ Link58.4Fast-iTPN-B (DINO 12ep, CLIP-distilled pre-training)2022-11-23
Mr. DETR: Instructive Multi-Route Training for Detection Transformers✓ Link58.476.363.940.862.875.3Mr. DETR (Swin-L, 12ep)2024-12-13
FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation✓ Link58.377.164.041.762.473.9H-Deformable-DETR + FeatAug-Flip (Swin-L, 24ep, 300 predictions)2023-03-02
MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism✓ Link58.276.563.442.562.874.6MI-DETR (Swin-L, 12ep)2025-03-03
Relation DETR: Exploring Explicit Position Relation Prior for Object Detection✓ Link58.176.463.541.863.073.5Relation-DETR (Swin-L, 24ep; official repo)2024-07-16
LP-DETR: Layer-wise Progressive Relations for Object Detection58.176.463.241.062.274.7LP-DETR (Swin-L, 12ep)2025-02-07
MDS-DETR: DETR with Masked Duplicate Suppressor✓ Link58.176.263.642.662.774.7MDS-DETR (Swin-L)2026-05-22
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows✓ Link58Swin-L (HTC++, multi scale)2021-03-25
Fractional Correspondence Framework in Detection Transformer57.976.163.641.562.374.9DINO + RTP (Swin-L, 12ep)2025-03-06
Relation DETR: Exploring Explicit Position Relation Prior for Object Detection✓ Link57.876.162.941.262.174.4Relation-DETR (Swin-L, 12ep)2024-07-16
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model57.841.561.273.9DI-MaskDINO (Swin-L, 50ep)2024-10-22
PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection57.876.262.641.462.274.2PaQ-DINO (Swin-L, 12ep)2026-03-06
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link57.7MOAT-1 (IN-1K pretraining, single-scale)2022-10-04
FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation✓ Link57.676.763.141.161.573.8Deformable-DETR w/ tricks + FeatAug-FC (Swin-L, 24ep, 300 predictions)2023-03-02
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models✓ Link57.675.863.241.461.774.3Frozen-DETR (Co-DINO-4scale, Swin-B IN-22K, frozen foundation model, 12ep)2024-10-25
Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders✓ Link57.676.463.240.861.873.6Dual-R-DETR (Swin-L, 12ep)2025-12-15
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality✓ Link57.4UM-MAE(HTC++, Swin-L, IN1K)2022-05-20
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement✓ Link57.375.562.340.961.874.5Salience DETR (FocalNet-L, 12ep; official repo)2024-03-24
YOLOv6 v3.0: A Full-Scale Reloading✓ Link57.274.5YOLOv6-L6(46 fps, 1280, V100)2023-01-13
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows✓ Link57.1Swin-L (HTC++, single scale)2021-03-25
TransNeXt: Robust Foveal Visual Perception for Vision Transformers✓ Link57.1TransNeXt-Base (IN-1K pretrain, DINO 1x)2023-11-28
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation✓ Link57.0Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale)2020-12-13
TransNeXt: Robust Foveal Visual Perception for Vision Transformers✓ Link56.6TransNeXt-Small (IN-1K pretrain, DINO 1x)2023-11-28
UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale✓ Link56.675.661.8254.8UniConvNet-L (IN-22K, Cascade Mask R-CNN 3x)2025-08-12
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement✓ Link56.575.061.540.261.272.8Salience DETR (Swin-L, 12ep; official repo)2024-03-24
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link56.4443UniRepLKNet-XL (IN-22K, Cascade Mask R-CNN 3x)2023-11-27
MogaNet: Multi-order Gated Aggregation Network✓ Link56.275.061.2238MogaNet-XL (IN-22K, Cascade Mask R-CNN 3x)2022-11-07
Instances as Queries✓ Link56.1QueryInst (Swin-L)2021-05-05
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection✓ Link56.1MViTv2-H (Cascade Mask R-CNN, single-scale, IN21k pre-train)2021-12-02
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link55.9MOAT-0 (IN-1K pretraining, single-scale)2022-10-04
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration✓ Link55.8Fast-iTPN-S (DINO 12ep, CLIP-distilled pre-training)2022-11-23
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link55.8276UniRepLKNet-L (IN-22K, Cascade Mask R-CNN 3x)2023-11-27
TransNeXt: Robust Foveal Visual Perception for Vision Transformers✓ Link55.7TransNeXt-Tiny (IN-1K pretrain, DINO 1x)2023-11-28
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models✓ Link55.773.961.338.458.872.3Frozen-DETR (DDQ + HPR, ResNet-50, frozen foundation model, 12ep)2024-10-25
UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale✓ Link55.774.460.4254.8UniConvNet-L (IN-22K, Cascade Mask R-CNN 1x)2025-08-12
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens55.474.159.9114SECViT-B (Cascade Mask R-CNN 3x)2024-05-22
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link55.2tiny-MOAT-3 (IN-1K pretraining, single-scale)2022-10-04
Understanding The Robustness in Vision Transformers✓ Link55.1FAN-L-Hybrid2022-04-26
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles✓ Link55Hiera-L2023-06-01
General Object Foundation Model for Images and Videos at Scale✓ Link55.0GLEE-Lite2023-12-14
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection✓ Link55.072.560.038.458.969.945DS-Det (Strip-MLP-T, 36ep)2025-07-26
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection✓ Link54.972.459.938.358.669.649DS-Det (Swin-T, 36ep)2025-07-26
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link54.8155UniRepLKNet-B (IN-22K, Cascade Mask R-CNN 3x)2023-11-27
Towards Sustainable Self-supervised Learning✓ Link54.6TEC(VIT-B, Mask-RCNN)2022-10-20
Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation✓ Link54.5Cascade Eff-B7 NAS-FPN (1280)2020-12-13
Context Autoencoder for Self-Supervised Representation Learning✓ Link54.575.260.1CAE* (ViT-L, 1600ep, Mask R-CNN 1x)2022-02-07
RMT: Retentive Networks Meet Vision Transformers✓ Link54.572.859.0111RMT-B (Cascade Mask R-CNN 3x)2023-09-20
Rotary Position Embedding for Vision Transformer✓ Link54.5Swin-B + RoPE-Mixed (DINO, 12ep)2024-03-20
Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection54.572.559.237.858.771.3BS-O2G (DEIM, ResNet-50, 60ep)2026-08-06
EfficientDet: Scalable and Efficient Object Detection✓ Link54.4EfficientDet-D7x (single-scale)2019-11-20
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection✓ Link54.3MViTv2-L (Cascade Mask R-CNN, single-scale)2021-12-02
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers✓ Link54.3319MixMAE (Swin-L, 600ep, Mask R-CNN)2022-05-26
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link54.3113UniRepLKNet-S (IN-22K, Cascade Mask R-CNN 3x)2023-11-27
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models✓ Link54.372.959.236.658.072.1Frozen-DETR (Co-DINO-4scale, ResNet-50, frozen foundation model, 24ep)2024-10-25
Rethinking Pre-training and Self-training✓ Link54.2SpineNet-190 (1280, with Self-training on OpenImages, single-scale)2020-06-11
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels53.9154OverLoCK-B (Cascade Mask R-CNN 3x)2025-02-27
Spectral-Adaptive Modulation Networks for Visual Perception✓ Link53.872.958.4SPANetV2-S36-hybrid (Cascade Mask R-CNN 3x)2025-03-31
Simple Training Strategies and Model Scaling for Object Detection✓ Link53.634.556.770.6Cascade RCNN-RS (SpineNet-143L, single scale)2021-06-30
Exploring Target Representations for Masked Autoencoders✓ Link53.6dBOT (ViT-B, CLIP teacher, Mask R-CNN)2022-09-08
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels53.6114OverLoCK-S (Cascade Mask R-CNN 3x)2025-02-27
Match One, Learn with Graph: One-to-Graph Query Collaboration with Backward Sharing for Object Detection53.571.358.136.958.070.5BS-O2G (DEIM, ResNet-50, 24ep)2026-08-06
EfficientDet: Scalable and Efficient Object Detection✓ Link53.4EfficientDet-D7 (1536)2019-11-20
MaxViT: Multi-Axis Vision Transformer✓ Link53.472.958.1157MaxViT-B (Cascade Mask R-CNN)2022-04-04
Masked Autoencoders Are Scalable Vision Learners✓ Link53.3MAE (ViT-L, Mask R-CNN)2021-11-11
MogaNet: Multi-order Gated Aggregation Network✓ Link53.371.857.8140MogaNet-L (Cascade Mask R-CNN 3x)2022-11-07
RMT: Retentive Networks Meet Vision Transformers✓ Link53.272.057.883RMT-S (Cascade Mask R-CNN 3x)2023-09-20
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models✓ Link53.271.858.035.156.570.6Frozen-DETR (DINO-4scale, ResNet-50, frozen foundation model, 24ep)2024-10-25
Simple Training Strategies and Model Scaling for Object Detection✓ Link53.133.956.270.3Cascade RCNN-RS (ResNet-200, single scale)2021-06-30
MaxViT: Multi-Axis Vision Transformer✓ Link53.172.558.1107MaxViT-S (Cascade Mask R-CNN)2022-04-04
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution53.1147PeLK-B-101 (Cascade Mask R-CNN 3x)2024-03-12
Spectral-Adaptive Modulation Networks for Visual Perception✓ Link53.172.257.6SPANetV2-S36-pure (Cascade Mask R-CNN 3x)2025-03-31
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link53.0tiny-MOAT-2 (IN-1K pretraining, single-scale)2022-10-04
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link53.0113UniRepLKNet-S (Cascade Mask R-CNN 3x)2023-11-27
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution52.9147PeLK-B (Cascade Mask R-CNN 3x)2024-03-12
Rotary Position Embedding for Vision Transformer✓ Link52.9ViT-L + RoPE-Mixed (DINO-ViTDet, 12ep)2024-03-20
DeepMIM: Deep Supervision for Masked Image Modeling✓ Link52.8DeepMIM-MAE-CLIP (ViT-B, Mask R-CNN)2023-03-15
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens52.873.657.775SECViT-B (Mask R-CNN 3x)2024-05-22
MambaVision: A Hybrid Mamba-Transformer Vision Backbone✓ Link52.871.357.2145MambaVision-B (Cascade Mask R-CNN 3x)2024-07-10
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection✓ Link52.7MViT-L (Mask R-CNN, single-scale, IN21k pre-train)2021-12-02
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers✓ Link52.7110MixMAE (Swin-B/W14, 600ep, Mask R-CNN)2022-05-26
MogaNet: Multi-order Gated Aggregation Network✓ Link52.672.057.3101MogaNet-B (Cascade Mask R-CNN 3x)2022-11-07
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs52.671.357.188.9HIRI-ViT-S (Cascade Mask R-CNN 3x, 1600px input)2024-03-18
PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection52.669.756.935.756.467.0PaQ-DINO (ResNet-50, 24ep)2026-03-06
Active Token Mixer✓ Link52.571.657.0ATMNet-B (Cascade Mask R-CNN 3x, MS)2022-03-11
LP-DETR: Layer-wise Progressive Relations for Object Detection52.570.057.236.256.367.1LP-DETR (ResNet-50, 24ep)2025-02-07
Language-aware Multiple Datasets Detection Pretraining for DETRs52.470.357.535.755.966.450METR (ResNet-50, Objects365+OpenImages pre-training, 12ep)2023-04-07
POA: Pre-training Once for Models of All Sizes✓ Link52.4POA ViT-B/16 (Cascade Mask R-CNN)2024-08-02
PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection52.469.857.436.155.666.2PaQ-DETR (ResNet-50, hybrid matching, 12ep)2026-03-06
MambaVision: A Hybrid Mamba-Transformer Vision Backbone✓ Link52.371.156.7108MambaVision-S (Cascade Mask R-CNN 3x)2024-07-10
LP-DETR: Layer-wise Progressive Relations for Object Detection52.369.656.835.855.966.6LP-DETR (ResNet-50, 12ep)2025-02-07
MDS-DETR: DETR with Masked Duplicate Suppressor✓ Link52.369.756.935.556.067.0MDS-DETR (ResNet-50, 900 queries, 24ep)2026-05-22
SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization✓ Link52.2Mask R-CNN (SpineNet-190, 1536x1536)2019-12-10
RMT: Retentive Networks Meet Vision Transformers✓ Link52.272.957.073RMT-B (Mask R-CNN 3x)2023-09-20
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution52.2108PeLK-S (Cascade Mask R-CNN 3x)2024-03-12
CoCAViT: Compact Vision Transformer with Robust Global Coordination52.271.056.8CoCAViT-28M (Cascade Mask R-CNN 3x)2025-08-07
MaxViT: Multi-Axis Vision Transformer✓ Link52.171.956.869MaxViT-T (Cascade Mask R-CNN)2022-04-04
Relation DETR: Exploring Explicit Position Relation Prior for Object Detection✓ Link52.169.756.636.156.066.5Relation-DETR (ResNet-50, 24ep)2024-07-16
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens52.073.557.3119SECViT-L (Mask R-CNN 1x)2024-05-22
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link51.9tiny-MOAT-1 (IN-1K pretraining, single-scale)2022-10-04
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs51.971.956.6113.0HIRI-ViT-S (Sparse R-CNN 3x, 1600px input)2024-03-18
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model51.936.354.766.7DI-MaskDINO (ResNet-50, 50ep)2024-10-22
PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection51.969.156.335.156.066.6PaQ-DINO (ResNet-50, 12ep)2026-03-06
Global Context Networks✓ Link51.870.456.1GCNet (ResNeXt-101 + DCN + cascade + GC r4)2020-12-24
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition✓ Link51.889UniRepLKNet-T (Cascade Mask R-CNN 3x)2023-11-27
LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking51.871.656.435.455.068.485AdaMixer + LORS (Swin-S, 3x)2024-03-07
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation51.870.456.734.654.467.5DINO + SACQ (Swin-T, 12ep)2024-05-06
CoCAViT: Compact Vision Transformer with Robust Global Coordination51.870.556.1CoCAViT-21M (Cascade Mask R-CNN 3x)2025-08-07
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss✓ Link51.769.056.335.555.066.1Align-DETR (ResNet-50, 24ep)2023-04-15
Relation DETR: Exploring Explicit Position Relation Prior for Object Detection✓ Link51.769.156.336.155.666.1Relation-DETR (ResNet-50, 12ep)2024-07-16
ELSA: Enhanced Local Self-Attention for Vision Transformer✓ Link51.670.556.0ELSA-S (Cascade Mask RCNN)2021-12-23
MogaNet: Multi-order Gated Aggregation Network✓ Link51.670.856.383MogaNet-S (Cascade Mask R-CNN 3x)2022-11-07
DeepMIM: Deep Supervision for Masked Image Modeling✓ Link51.6DeepMIM-MAE (ViT-B, Mask R-CNN)2023-03-15
RMT: Retentive Networks Meet Vision Transformers✓ Link51.673.156.5114RMT-L (Mask R-CNN 1x)2023-09-20
Focal Modulation Networks✓ Link51.570.356.0FocalNet-T (LRF, Cascade Mask R-CNN)2022-03-22
SpectFormer: Frequency and Attention is what you need in a Vision Transformer51.570.256.3SpectFormer-H-S (Cascade Mask R-CNN 3x)2023-04-13
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution51.486PeLK-T (Cascade Mask R-CNN 3x)2024-03-12
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels51.4114OverLoCK-B (Mask R-CNN 3x)2025-02-27
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection✓ Link51.369.15634.554.265.8DINO-5scale (24 epoch)2022-03-07
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection✓ Link51.368.455.935.554.865.748DS-Det (ResNet-50, 24ep)2025-07-26
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection✓ Link51.26955.83554.365.3DINO-5scale (36 epoch)2022-03-07
Rotary Position Embedding for Vision Transformer✓ Link51.2ViT-B + RoPE-Mixed (DINO-ViTDet, 12ep)2024-03-20
UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale✓ Link51.272.256.1118.0UniConvNet-B (Mask R-CNN 3x)2025-08-12
MDS-DETR: DETR with Masked Duplicate Suppressor✓ Link51.268.955.834.054.866.2Hybrid-MDS-DETR (ResNet-50, 12ep)2026-05-22
MambaVision: A Hybrid Mamba-Transformer Vision Backbone✓ Link51.170.055.686MambaVision-T (Cascade Mask R-CNN 3x)2024-07-10
MDS-DETR: DETR with Masked Duplicate Suppressor✓ Link51.168.855.534.055.066.0MDS-DETR (ResNet-50, 900 queries, 12ep)2026-05-22
ResNeSt: Split-Attention Networks✓ Link50.91ResNeSt-200 (Cascade R-CNN, DCN, tricks, 3x; official repo)2020-04-19
RF-Next: Efficient Receptive Field Search for Convolutional Neural Networks✓ Link50.969.555.534.354.665.8RF-ConvNeXt-T (Cascade Mask R-CNN)2022-06-14
Learning Data Augmentation Strategies for Object Detection✓ Link50.734.255.564.5NAS-FPN (AmoebaNet-D, learned aug)2019-06-26
Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders✓ Link50.768.855.034.154.264.6Dual-R-DETR (ResNet-50, 12ep)2025-12-15
POA: Pre-training Once for Models of All Sizes✓ Link50.6POA ViT-S/16 (Cascade Mask R-CNN)2024-08-02
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer50.671.355.3140Hyb-KAN ViT-B (Mask R-CNN, ViTDet-style, 3x)2025-05-07
ResNeSt: Split-Attention Networks✓ Link50.54ResNeSt-200 (Cascade R-CNN, tricks, 3x; official repo)2020-04-19
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models✓ Link50.5tiny-MOAT-0 (IN-1K pretraining, single-scale)2022-10-04
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss✓ Link50.567.755.334.753.664.6Align-DETR (ResNet-50, 12ep)2023-04-15
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection✓ Link50.567.555.134.354.664.548DS-Det (ResNet-50, 12ep)2025-07-26
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems50.5VisionHOPE-B (Mask R-CNN 1x)2026-09-27
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems50.571.755.5VisionHOPE-S (Mask R-CNN 3x)2026-09-27
FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation✓ Link50.468.754.932.753.665.2H-Deformable-DETR + FeatAug-Flip (ResNet-50, 24ep)2023-03-02
Fractional Correspondence Framework in Detection Transformer50.467.955.234.753.864.2RTP-DETR (ResNet-50, 12ep)2025-03-06
Masked Autoencoders Are Scalable Vision Learners✓ Link50.3MAE (ViT-B, Mask R-CNN)2021-11-11
SpectFormer: Frequency and Attention is what you need in a Vision Transformer50.370.055.2SpectFormer-H-S (GFL 3x)2023-04-13
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention✓ Link50.271.554.034.754.665.363DAT-S++ (RetinaNet 3x)2023-09-04
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens50.271.453.933.254.566.3105SECViT-L (RetinaNet 1x)2024-05-22
PVT v2: Improved Baselines with Pyramid Vision Transformer✓ Link50.169.554.9Sparse R-CNN (PVTv2-B2)2021-06-25
GlobalMamba: Global Image Serialization for Vision Mamba✓ Link50.154.970GlobalMamba-S (Mask R-CNN 3x)2024-10-14
UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale✓ Link50.171.054.850.0UniConvNet-T (Mask R-CNN 3x)2025-08-12
Pix2seq: A Language Modeling Framework for Object Detection✓ Link50.0Pix2seq (ViT-L)2021-09-22
V2M: Visual 2-Dimensional Mamba for Image Representation Learning✓ Link50.070.954.870V2M-B (Mask R-CNN 3x)2024-10-14
DaViT: Dual Attention Vision Transformers✓ Link49.9DaViT-T (Mask R-CNN, 36 epochs)2022-04-07
FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation✓ Link49.968.354.732.552.965.1Deformable-DETR w/ tricks + FeatAug-FC (ResNet-50, 24ep)2023-03-02
VMamba: Visual State Space Model✓ Link49.970.954.770VMamba-S (Mask R-CNN 3x)2024-01-18
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels49.9114OverLoCK-B (Mask R-CNN 1x)2025-02-27
Enhanced Training of Query-Based Object Detection via Selective Query Recollection✓ Link49.868.854.032.053.465.1SQR-AdaMixer (ResNet-101)2022-12-15
Bottleneck Transformers for Visual Recognition49.771.354.6BoTNet 200 (Mask R-CNN, single scale, 72 epochs)2021-01-27
Dynamic Head: Unifying Object Detection Heads with Attentions✓ Link49.768.054.333.354.264.2DyHead (Swin-T, 2x)2021-06-15
Language-aware Multiple Datasets Detection Pretraining for DETRs49.767.653.832.252.664.950METR (ResNet-50, 50ep)2023-04-07
Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics✓ Link49.771.354.6VNCT-T (Mask R-CNN 3x)2026-07-03
Spectral-Adaptive Modulation Networks for Visual Perception✓ Link49.671.354.8SPANetV2-S18-hybrid (Mask R-CNN 3x)2025-03-31
Bottleneck Transformers for Visual Recognition49.57154.2BoTNet 152 (Mask R-CNN, single scale, 72 epochs)2021-01-27
DN-DETR: Accelerate DETR Training by Introducing Query DeNoising✓ Link49.567.653.831.352.665.447DN-Deformable-DETR-R50++2022-03-02
Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN✓ Link49.4A2MIM (ViT-B, Mask R-CNN)2022-05-27
MogaNet: Multi-order Gated Aggregation Network✓ Link49.470.754.1102MogaNet-L (Mask R-CNN 1x)2022-11-07
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation49.466.953.931.852.364.658DINO + SACQ (ResNet-50, 12ep)2024-05-06
Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders✓ Link49.467.953.933.152.664.0Deformable-DETR++ + Dual-R (ResNet-50, 24ep)2025-12-15
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems49.470.654.5VisionHOPE-T (Mask R-CNN 3x)2026-09-27
GlobalMamba: Global Image Serialization for Vision Mamba✓ Link49.371.454.2108GlobalMamba-B (Mask R-CNN 1x)2024-10-14
ViDT: An Efficient and Effective Fully Transformer-based Object Detector✓ Link49.269.453.130.652.666.9ViDT (Swin-base, 50ep)2021-10-08
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention✓ Link49.270.353.032.753.464.734DAT-T++ (RetinaNet 3x)2023-09-04
VMamba: Visual State Space Model✓ Link49.271.454.0108VMamba-B (Mask R-CNN 1x)2024-01-18
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement✓ Link49.267.153.832.753.063.156.1Salience DETR (ResNet-50, 12ep)2024-03-24
Recurrent Glimpse-based Decoder for Detection with Transformer✓ Link49.167.553.13052.665REGO-Deformable DETR-X1012021-12-09
Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion✓ Link49.167.253.230.552.664.755SAM-DETR++ w/ SMCA+DN (ResNet-50, 50ep)2022-07-28
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs49.171.054.369.4HIRI-ViT-B (Mask R-CNN 1x, 1600px input)2024-03-18
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm✓ Link49.070.353.668EATFormer-Base (Mask R-CNN 3x)2022-06-19
V2M: Visual 2-Dimensional Mamba for Image Representation Learning✓ Link49.070.653.550V2M-S (Mask R-CNN 3x)2024-10-14
GlobalMamba: Global Image Serialization for Vision Mamba✓ Link49.070.553.750GlobalMamba-T (Mask R-CNN 3x)2024-10-14
GlobalMamba: Global Image Serialization for Vision Mamba✓ Link49.070.553.570GlobalMamba-S (Mask R-CNN 1x)2024-10-14
Enhanced Training of Query-Based Object Detection via Selective Query Recollection✓ Link48.967.553.232.051.863.7SQR-AdaMixer (ResNet-50)2022-12-15
V2M: Visual 2-Dimensional Mamba for Image Representation Learning✓ Link48.970.253.670V2M-B (Mask R-CNN 1x)2024-10-14
VMamba: Visual State Space Model✓ Link48.870.453.550VMamba-T (Mask R-CNN 3x)2024-01-18
MogaNet: Multi-order Gated Aggregation Network✓ Link48.769.552.631.553.462.792MogaNet-L (RetinaNet 1x)2022-11-07
VMamba: Visual State Space Model✓ Link48.770.053.470VMamba-S (Mask R-CNN 1x)2024-01-18
Rethinking ImageNet Pre-training48.666.852.9Mask R-CNN (ResNeXt-152-FPN, cascade)2018-11-21
CenterMask : Real-Time Anchor-Free Instance Segmentation✓ Link48.6CenterMask2 (VoVNetV2-99, 3x, TTA; official repo)2019-11-15
USB: Universal-Scale Object Detection Benchmark✓ Link48.667.152.730.153.063.8UniverseNet-20.08d (Res2Net-50, DCN)2021-03-25
BiFormer: Vision Transformer with Bi-Level Routing Attention✓ Link48.670.553.8BiFormer-B (Mask R-CNN 1x)2023-03-15
XCiT: Cross-Covariance Image Transformers✓ Link48.5XCiT-M24/82021-06-17
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention✓ Link48.570.253.3DeBiFormer-B (Mask R-CNN 1x)2024-10-11
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs48.469.751.933.452.663.459.8HIRI-ViT-B (RetinaNet 1x, 1600px input)2024-03-18
ELSA: Enhanced Local Self-Attention for Vision Transformer✓ Link48.370.452.9ELSA-S (Mask RCNN)2021-12-23
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer48.369.852.751.7Hyb-KAN ViT-S (Mask R-CNN, ViTDet-style, 3x)2025-05-07
LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking48.267.552.631.751.363.879AdaMixer + LORS (ResNet-101, 3x)2024-03-07
XCiT: Cross-Covariance Image Transformers✓ Link48.1XCiT-S24/82021-06-17
GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond✓ Link47.966.952.2GCNet (ResNeXt-101 + DCN + cascade + GC r16)2019-04-25
Vision Transformer with Deformable Attention✓ Link47.969.651.232.351.863.4DAT-S (RetinaNet 3x)2022-01-03
MogaNet: Multi-order Gated Aggregation Network✓ Link47.970.052.763MogaNet-B (Mask R-CNN 1x)2022-11-07
BiFormer: Vision Transformer with Bi-Level Routing Attention✓ Link47.869.852.3BiFormer-S (Mask R-CNN 1x)2023-03-15
Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics✓ Link47.870.352.6VNCT-T (Mask R-CNN 1x)2026-07-03
MogaNet: Multi-order Gated Aggregation Network✓ Link47.768.951.030.552.261.754MogaNet-B (RetinaNet 1x)2022-11-07
MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection✓ Link47.630.251.860.8MAE-DET-L + GFLV22021-11-26
Language-aware Multiple Datasets Detection Pretraining for DETRs47.665.652.029.850.962.650METR (ResNet-50, 12ep)2023-04-07
LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking47.666.652.031.150.262.560AdaMixer + LORS (ResNet-50, 3x)2024-03-07
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation47.665.751.930.050.562.458DAB-Deformable-DETR + SACQ (ResNet-50, 50ep)2024-05-06
ViDT: An Efficient and Effective Fully Transformer-based Object Detector✓ Link47.567.751.429.250.764.861ViDT (Swin-small, 50ep)2021-10-08
Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion✓ Link47.566.551.329.350.862.755SAM-DETR++ (ResNet-50, 50ep)2022-07-28
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention✓ Link47.569.752.1DeBiFormer-S (Mask R-CNN 1x)2024-10-11
Rethinking ImageNet Pre-training47.4Mask R-CNN (ResNet-101-FPN, GN, Cascade)2018-11-21
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm✓ Link47.469.351.944EATFormer-Small (Mask R-CNN 3x)2022-06-19
MambaOut: Do We Really Need Mamba for Vision?✓ Link47.469.152.465MambaOut-Small (Mask R-CNN 1x)2024-05-13
MambaOut: Do We Really Need Mamba for Vision?✓ Link47.469.352.2100MambaOut-Base (Mask R-CNN 1x)2024-05-13
Pix2seq: A Language Modeling Framework for Object Detection✓ Link47.3Pix2seq (R50-C4)2021-09-22
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation47.366.751.230.650.062.651Two-stage Deformable-DETR + SACQ (ResNet-50, 50ep)2024-05-06
EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm✓ Link47.269.452.168EATFormer-Base (Mask R-CNN 1x)2022-06-19
Pix2seq: A Language Modeling Framework for Object Detection✓ Link47.1Pix2seq (ViT-B)2021-09-22
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention✓ Link47.168.250.230.351.163.0DeBiFormer-B (RetinaNet 1x)2024-10-11
Deep High-Resolution Representation Learning for Visual Recognition✓ Link47.028.850.362.2HTC (HRNetV2p-W48)2019-08-20
Augmenting Convolutional networks with attention-based aggregation✓ Link47.0PatchConvNet-S120 (Mask R-CNN)2021-12-27
SpectFormer: Frequency and Attention is what you need in a Vision Transformer46.968.851.8SpectFormer-B (Mask R-CNN)2023-04-13
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model46.928.849.562.9DI-MaskDINO (ResNet-50, 12ep)2024-10-22
RepPoints: Point Set Representation for Object Detection✓ Link46.8RPDet (ResNeXt-101-DCN, multi-scale)2019-04-25
Cross Resolution Encoding-Decoding For Detection Transformers✓ Link46.866.850.527.450.764.045DN-DETR + CRED-OO (ResNet-50, 50ep)2024-10-05
MogaNet: Multi-order Gated Aggregation Network✓ Link46.768.051.345MogaNet-S (Mask R-CNN 1x)2022-11-07
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR✓ Link46.66750.228.150.564.163DAB-DETR-DC5-R1012022-01-28
LBMamba: Locally Bi-directional Mamba✓ Link46.665.250.527.050.663.674LBVim-300 (Cascade Mask R-CNN)2025-06-19
Rethinking ImageNet Pre-training46.467.151.1Mask R-CNN (ResNeXt-152-FPN)2018-11-21
RepPoints: Point Set Representation for Object Detection✓ Link46.4RPDet (ResNet-101-DCN, multi-scale)2019-04-25
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals✓ Link46.464.649.528.348.361.6Sparse R-CNN (ResNet-101, learnable proposals, random crop aug, FPN)2020-11-25
Augmenting Convolutional networks with attention-based aggregation✓ Link46.4PatchConvNet-S60 (Mask R-CNN)2021-12-27
AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs46.468.551.332.3AttentionViG-B (Mask R-CNN 1x)2025-09-29
PVT v2: Improved Baselines with Pyramid Vision Transformer✓ Link46.364.350.5Cascade Mask R-CNN (ResNet-50, 3x; reported by PVTv2)2021-06-25
Cross Resolution Encoding-Decoding For Detection Transformers✓ Link46.265.849.826.850.063.545DN-DETR + CRED (ResNet-50, 50ep)2024-10-05
HoughNet: Integrating near and long-range evidence for bottom-up object detection✓ Link46.164.650.330.048.859.7HoughNet (HG-104, MS)2020-07-05
Deep High-Resolution Representation Learning for Visual Recognition✓ Link46.027.548.960.1Cascade Mask R-CNN (HRNetV2p-W48)2019-08-20
SpectFormer: Frequency and Attention is what you need in a Vision Transformer46.066.449.729.549.761.1SpectFormer-B (RetinaNet)2023-04-13
Bottleneck Transformers for Visual Recognition45.9BoTNet 50 (72 epochs)2021-01-27
Conditional DETR for Fast Training Convergence✓ Link45.966.849.527.250.363.363Conditional DETR-DC5-R1012021-08-13
MogaNet: Multi-order Gated Aggregation Network✓ Link45.866.649.029.150.159.835MogaNet-S (RetinaNet 1x)2022-11-07
CenterMask : Real-Time Anchor-Free Instance Segmentation✓ Link45.629.249.358.8CenterMask (VoVNetV2-99, 3x; official repo)2019-11-15
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention✓ Link45.666.648.928.749.361.6DeBiFormer-S (RetinaNet 1x)2024-10-11
Cross Resolution Encoding-Decoding For Detection Transformers✓ Link45.464.949.427.048.562.245DAB-DETR + CRED (ResNet-50, 50ep)2024-10-05
LBMamba: Locally Bi-directional Mamba✓ Link45.463.849.325.549.462.466LBVim-Ti (Cascade Mask R-CNN)2025-06-19
Deep High-Resolution Representation Learning for Visual Recognition✓ Link45.327.048.459.5HTC (HRNetV2p-W32)2019-08-20
Conditional DETR for Fast Training Convergence✓ Link45.165.448.525.34962.244Conditional DETR-DC5-R502021-08-13
Anchor DETR: Query Design for Transformer-Based Object Detection✓ Link45.165.748.825.849.461.6Anchor DETR-DC5-R1012021-09-15
MambaOut: Do We Really Need Mamba for Vision?✓ Link45.167.349.643MambaOut-Tiny (Mask R-CNN 1x)2024-05-13
Non-local Neural Networks✓ Link45.067.848.9Mask R-CNN (ResNeXt-152 + 1 NL)2017-11-21
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals✓ Link45.063.448.226.947.259.5Sparse R-CNN (ResNet-50, learnable proposals, random crop aug, FPN)2020-11-25
Pix2seq: A Language Modeling Framework for Object Detection✓ Link45.063.248.628.248.960.4Pix2seq (R101-DC5)2021-09-22
Attentive Normalization✓ Link44.966.249.1Mask R-CNN-FPN (AOGNet-40M)2019-08-04
CenterMask : Real-Time Anchor-Free Instance Segmentation✓ Link44.9Mask R-CNN (VoVNetV2-99, 3x; CenterMask2 repo)2019-11-15
End-to-End Object Detection with Transformers✓ Link44.964.747.723.749.562.3DETR-DC5 (ResNet-101)2020-05-26
RepPoints: Point Set Representation for Object Detection✓ Link44.8RPDet (ResNet-101-DCN, multi-scale train)2019-04-25
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing✓ Link44.864.348.926.648.359.6R3-CNN (ResNet-50-FPN, DCN)2021-04-03
ViDT: An Efficient and Effective Fully Transformer-based Object Detector✓ Link44.864.548.725.947.662.138ViDT (Swin-tiny, 50ep)2021-10-08
Semantic-Aligned Matching for Enhanced DETR Convergence and Multi-Scale Feature Fusion✓ Link44.862.647.926.748.260.955SAM-DETR++ w/ SMCA+DN (ResNet-50, 12ep)2022-07-28
Deep High-Resolution Representation Learning for Visual Recognition✓ Link44.662.748.726.348.158.5Cascade R-CNN (HRNetV2p-W48)2019-08-20
CenterMask : Real-Time Anchor-Free Instance Segmentation✓ Link44.627.748.357.3CenterMask (VoVNetV2-57, 3x; official repo)2019-11-15
RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations✓ Link44.666.349.0RecNeXt-M5 (Mask R-CNN)2024-12-27
RepPoints: Point Set Representation for Object Detection✓ Link44.5RPDet (ResNeXt-101-DCN)2019-04-25
Deep High-Resolution Representation Learning for Visual Recognition✓ Link44.526.147.958.5Cascade Mask R-CNN (HRNetV2p-W32)2019-08-20
PVT v2: Improved Baselines with Pyramid Vision Transformer✓ Link44.563.048.3GFL (ResNet-50, 3x; reported by PVTv2)2021-06-25
Conditional DETR for Fast Training Convergence✓ Link44.565.647.523.648.463.663Conditional DETR-R1012021-08-13
CenterMask : Real-Time Anchor-Free Instance Segmentation✓ Link44.426.747.757.1CenterMask (X-101-32x8d, 3x; official repo)2019-11-15
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing✓ Link44.364.148.42747.158.9R3-CNN (ResNet-50-FPN, GC-Net)2021-04-03
Anchor DETR: Query Design for Transformer-Based Object Detection✓ Link44.264.747.524.748.260.6Anchor DETR-DC5-R502021-09-15
SCSA: Exploring the Synergistic Effects Between Spatial and Channel Attention44.263.148.226.048.257.588.52Cascade R-CNN + SCSA (ResNet-101, 1x)2024-07-06
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals✓ Link44.162.147.226.146.359.7Sparse R-CNN (ResNet-101, FPN)2020-11-25
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR✓ Link44.164.747.224.148.262.963DAB-DETR-R1012022-01-28
Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond44.0764.4848.4026.7348.1156.80Faster R-CNN + LossMix (ResNet-101-FPN, 3x)2023-03-18
End-to-End Object Detection with Transformers✓ Link4463.947.827.248.156Faster RCNN-R101-FPN+2020-05-26
Deep High-Resolution Representation Learning for Visual Recognition✓ Link43.761.747.725.646.557.4Cascade R-CNN (HRNetV2p-W32)2019-08-20
Micro-Batch Training with Batch-Channel Normalization and Weight Standardization✓ Link43.664.447.925.647.557.4Mask R-CNN-FPN (ResNet-101, GN+WS+BCN)2019-03-25
PVT v2: Improved Baselines with Pyramid Vision Transformer✓ Link43.561.947.0ATSS (ResNet-50, 3x; reported by PVTv2)2021-06-25
RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations✓ Link43.564.947.7RecNeXt-M4 (Mask R-CNN)2024-12-27
AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs43.565.847.612.3AttentionViG-S (Mask R-CNN 1x)2025-09-29
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions✓ Link43.463.646.126.146.059.5PVT-Large (RetinaNet 3x,MS)2021-02-24
Bottom-up Object Detection by Grouping Extreme and Center Points✓ Link43.359.646.825.746.659.4ExtremeNet (Hourglass-104, multi-scale)2019-01-23
Hybrid Task Cascade for Instance Segmentation✓ Link43.259.440.720.340.952.3HTC (cascade)2019-01-22
Pix2seq: A Language Modeling Framework for Object Detection✓ Link43.261.046.126.64758.6Pix2seq (R50-DC5 )2021-09-22
Deformable ConvNets v2: More Deformable, Better Results43.1Mask R-CNN (ResNet-101, DCNv2)2018-11-27
Deep High-Resolution Representation Learning for Visual Recognition✓ Link43.126.646.056.9HTC (HRNetV2p-W18)2019-08-20
HoughNet: Integrating near and long-range evidence for bottom-up object detection✓ Link43.062.246.925.547.655.8HoughNet (HG-104)2020-07-05
Conditional DETR for Fast Training Convergence✓ Link436445.722.746.761.544Conditional DETR-R502021-08-13
Scaling Graph Convolutions for Mobile Vision✓ Link43.064.947.127.7MobileViGv2-B (Mask R-CNN 1x)2024-06-09
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals✓ Link42.861.245.726.744.657.6Sparse R-CNN (ResNet-50, FPN)2020-11-25
X-volution: On the unification of convolution and self-attention42.86446.426.94655Faster R-CNN (FPN, X-volution)2021-06-04
UniHead: Unifying Multi-Perception for Detection Heads✓ Link42.861.046.224.846.856.8UniHead (ResNet-50, 1x)2023-09-23
Cascade R-CNN: Delving into High Quality Object Detection✓ Link42.761.646.623.846.257.4Cascade R-CNN (ResNet-101-FPN+, cascade)2017-12-03
Dynamic Feature Pyramid Networks for Object Detection42.7Cascade R-CNN + DyFPN (ResNet-101, 1x)2020-12-01
CornerNet-Lite: Efficient Keypoint Based Object Detection✓ Link42.625.544.358.4CornerNet-Saccade (Hourglass-54)2019-04-18
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions✓ Link42.663.745.425.846.058.4PVT-Large (RetinaNet 1x)2021-02-24
Pix2seq: A Language Modeling Framework for Object Detection✓ Link42.6Pix2seq (R50)2021-09-22
MogaNet: Multi-order Gated Aggregation Network✓ Link42.664.046.425MogaNet-T (Mask R-CNN 1x)2022-11-07
S2AFormer: Strip Self-Attention for Efficient Vision Transformer42.664.546.942S2AFormer-M (Mask R-CNN 1x)2025-05-28
Scaling Graph Convolutions for Mobile Vision✓ Link42.563.946.315.4MobileViGv2-M (Mask R-CNN 1x)2024-06-09
Group Normalization✓ Link42.362.846.2Mask R-CNN (ResNet-101-FPN, GroupNorm, long)2018-03-22
Deep High-Resolution Representation Learning for Visual Recognition✓ Link42.325.045.454.9Mask R-CNN (HRNetV2p-W32)2019-08-20
Rethinking and Improving Relative Position Encoding for Vision Transformer✓ Link42.3DETR-ResNet50 with iRPE-K (300 epochs)2021-07-29
UniHead: Unifying Multi-Perception for Detection Heads✓ Link42.361.145.624.546.055.331.81GFL + UniHead (ResNet-50, 1x)2023-09-23
Scale-Aware Trident Networks for Object Detection✓ Link4263.545.524.94756.9TridentNet (ResNet-101)2019-01-07
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing✓ Link426146.324.545.255.7R3-CNN (ResNet-50-FPN)2021-04-03
Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond41.8262.5145.8125.0445.4854.03Faster R-CNN + LossMix (ResNet-50-FPN, 3x)2023-03-18
Deep High-Resolution Representation Learning for Visual Recognition✓ Link41.862.845.944.754.6Faster R-CNN (HRNetV2p-W48)2019-08-20
Deformable ConvNets v2: More Deformable, Better Results41.722.245.858.7Faster R-CNN (ResNet-101, DCNv2)2018-11-27
LIP: Local Importance-based Pooling✓ Link41.763.645.625.245.8Faster R-CNN (LIP-ResNet-101)2019-08-12
RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations✓ Link41.763.445.4RecNeXt-M3 (Mask R-CNN)2024-12-27
S2AFormer: Strip Self-Attention for Efficient Vision Transformer41.762.444.525.844.655.432S2AFormer-M (RetinaNet 1x)2025-05-28
Feature Selective Anchor-Free Module for Single-Shot Object Detection41.662.4FSAF (ResNeXt-101, anchor-based branches)2019-03-02
SCSA: Exploring the Synergistic Effects Between Spatial and Channel Attention41.562.945.424.645.353.760.88Faster R-CNN + SCSA (ResNet-101, 1x)2024-07-06
CornerNet-Lite: Efficient Keypoint Based Object Detection✓ Link41.423.843.557.1CornerNet-Saccade (Hourglass-104)2019-04-18
MogaNet: Multi-order Gated Aggregation Network✓ Link41.461.544.425.145.753.614MogaNet-T (RetinaNet 1x)2022-11-07
Grid R-CNN✓ Link41.360.344.423.445.854.1Grid R-CNN (ResNet-101-FPN)2018-11-29
CenterNet: Keypoint Triplets for Object Detection✓ Link41.359.243.923.643.855.8CenterNet511 (Hourglass-52)2019-04-17
Deep High-Resolution Representation Learning for Visual Recognition✓ Link41.359.244.923.744.254.1Cascade R-CNN (HRNetV2p-W18)2019-08-20
RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free✓ Link41.160.244.1RetinaMask (ResNet-101-FPN)2019-01-10
MetaFormer Is Actually What You Need for Vision✓ Link41.063.144.8PoolFormer-S36 (Mask R-CNN)2021-11-22
LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection✓ Link41.057.944.321.946.156.82.4LeYOLO-Large (768)2024-06-20
Deep High-Resolution Representation Learning for Visual Recognition✓ Link40.961.844.824.443.753.3Faster R-CNN (HRNetV2p-W32)2019-08-20
VirTex: Learning Visual Representations from Textual Annotations✓ Link40.9VirTex Mask R-CNN (ResNet-50-FPN)2020-06-11
Dynamic Feature Pyramid Networks for Object Detection40.9Mask R-CNN + DyFPN (ResNet-101, 1x)2020-12-01
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices✓ Link40.957.63.30PP-PicoDet-L (640)2021-11-01
Non-local Neural Networks✓ Link40.863.144.5Mask R-CNN (ResNet-101 + 1 NL)2017-11-21
Group Normalization✓ Link40.861.644.4Mask R-CNN (ResNet-50-FPN, GroupNorm, long)2018-03-22
RepPoints: Point Set Representation for Object Detection✓ Link40.8RPDet (ResNet-50, multi-scale train)2019-04-25
Rethinking and Improving Relative Position Encoding for Vision Transformer✓ Link40.8DETR-ResNet50 with iRPE-K (150 epochs)2021-07-29
LSNet: See Large, Focus Small✓ Link40.863.444.0LSNet-B (Mask R-CNN 1x)2025-03-29
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection✓ Link40.760.743.3Faster R-CNN+aLRP Loss (ResNet-50, 500 scale)2020-09-28
MogaNet: Multi-order Gated Aggregation Network✓ Link40.762.344.423MogaNet-XT (Mask R-CNN 1x)2022-11-07
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding40.763.344.242.2ParFormer-M (Mask R-CNN 1x)2024-03-22
Acquisition of Localization Confidence for Accurate Object Detection✓ Link40.659.0IoU-Net (ResNet-101-FPN, IoU-NMS + refinement)2018-07-30
Reducing Label Noise in Anchor-Free Object Detection✓ Link40.559.544.225.444.752.3PPDet (ResNet-101-FPN)2020-08-03
ViDT: An Efficient and Effective Fully Transformer-based Object Detector✓ Link40.459.643.323.242.555.816ViDT (Swin-nano, 50ep)2021-10-08
A lightweight mechanism for vision-transformer-based object detection40.458.943.924.043.952.0XFCOS (ResNet-50, 3x)2025-05-22
Cascade R-CNN: Delving into High Quality Object Detection✓ Link40.359.443.722.943.754.1Cascade R-CNN (ResNet-50-FPN+)2017-12-03
Group Normalization✓ Link40.36144Mask R-CNN (ResNet-50-FPN, GroupNorm)2018-03-22
Bottom-up Object Detection by Grouping Extreme and Center Points✓ Link40.355.143.721.644.056.1ExtremeNet (Hourglass-104, single-scale)2019-01-23
RepPoints: Point Set Representation for Object Detection✓ Link40.3RPDet (ResNet-101)2019-04-25
A novel Region of Interest Extraction Layer for Instance Segmentation✓ Link40.362.44424.244.452.5GC-Net + GRoIE (ResNet-50-FPN)2020-04-28
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection✓ Link40.260.342.3RetinaNet+aLRP Loss (ResNet-50, 500 scale)2020-09-28
Cross-Iteration Batch Normalization✓ Link40.160.544.1Mask R-CNN (ResNet-101-FPN, CBN)2020-02-13
Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN✓ Link39.8A2MIM (ResNet-50-C4, Mask R-CNN 2x)2022-05-27
A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection✓ Link39.758.841.5FoveaBox+aLRP Loss (ResNet-50, 500 scale)2020-09-28
MogaNet: Multi-order Gated Aggregation Network✓ Link39.760.042.423.843.651.712MogaNet-XT (RetinaNet 1x)2022-11-07
Grid R-CNN✓ Link39.658.342.422.643.851.5Grid R-CNN (ResNet-50-FPN)2018-11-29
Adaptively Connected Neural Networks✓ Link39.5Mask R-CNN (ResNet-50, ACNet)2019-04-07
Feature Selective Anchor-Free Module for Single-Shot Object Detection39.359.2FSAF (ResNet-101, anchor-based branches)2019-03-02
LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection✓ Link39.355.742.518.844.156.1LeYOLO-Medium (640)2024-06-20
Deep High-Resolution Representation Learning for Visual Recognition✓ Link39.223.741.751.0Mask R-CNN (HRNetV2p-W18)2019-08-20
LSNet: See Large, Focus Small✓ Link39.260.041.522.143.052.9LSNet-B (RetinaNet 1x)2025-03-29
Non-local Neural Networks✓ Link39.061.141.9Mask R-CNN (ResNet-50 + 1 NL)2017-11-21
FoveaBox: Beyond Anchor-based Object Detector✓ Link38.958.441.522.343.551.7FoveaBox (ResNet-101-FPN, 800x800)2019-04-08
FCOS: Fully Convolutional One-Stage Object Detection✓ Link38.657.441.422.342.549.8FCOS (ResNet-50-FPN + improvements)2019-04-02
RepPoints: Point Set Representation for Object Detection✓ Link38.6RPDet (ResNet-50)2019-04-25
Libra R-CNN: Towards Balanced Learning for Object Detection✓ Link38.559.342.022.942.150.5Libra R-CNN (ResNet-50 FPN)2019-04-04
CornerNet: Detecting Objects as Paired Keypoints✓ Link38.453.840.918.640.551.8CornerNet511 (Hourglass-104)2018-08-03
A novel Region of Interest Extraction Layer for Instance Segmentation✓ Link38.459.941.722.942.149.7Mask R-CNN (ResNet-50-FPN, GRoIE)2020-04-28
LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection✓ Link38.254.141.317.642.255.11.9LeYOLO-Small (640)2024-06-20
FoveaBox: Beyond Anchor-based Object Detector✓ Link3857.840.219.542.252.7FoveaBox (ResNet-101-FPN, 600x600)2019-04-08
Deep High-Resolution Representation Learning for Visual Recognition✓ Link38.058.941.522.640.849.6Faster R-CNN (HRNetV2p-W18)2019-08-20
Feature Selective Anchor-Free Module for Single-Shot Object Detection37.958.0FSAF (ResNet-101)2019-03-02
ELA: Efficient Local Attention for Deep Convolutional Neural Networks37.8157.940.218.842.652.4YOLOF + ELA-L (1x)2024-03-02
A novel Region of Interest Extraction Layer for Instance Segmentation✓ Link37.559.240.622.341.547.8Faster R-CNN (ResNet-50-FPN, GRoIE)2020-04-28
Boosting Convolutional Neural Networks with Middle Spectrum Grouped Convolution✓ Link37.459.140.2Faster R-CNN (ResNet-50-MSGC, 1x)2023-04-13
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding37.359.040.425.7ParFormer-T (Mask R-CNN 1x)2024-03-22
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices✓ Link37.256.07.2YOLOv5s (640, reported by PP-PicoDet)2021-11-01
torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation✓ Link36.9Mask R-CNN (Bottleneck-injected ResNet-50, FPN)2020-11-25
Mask R-CNN✓ Link36.759.538.9Mask R-CNN (ResNeXt-101-FPN)2017-03-20
FoveaBox: Beyond Anchor-based Object Detector✓ Link36.055.237.918.639.450.5FoveaBox (ResNet-50-FPN, 600x600)2019-04-08
Feature Selective Anchor-Free Module for Single-Shot Object Detection35.955.037.919.839.648.2FSAF (ResNet-50)2019-03-02
torchdistill: A Modular, Configuration-Driven Framework for Knowledge Distillation✓ Link35.9Faster R-CNN (Bottleneck-injected ResNet-50 and FPN)2020-11-25
Gradient Harmonized Single-stage Detector✓ Link35.855.538.119.639.646.7GHM-C + GHM-R (RetinaNet-FPN-ResNet-50, M=30)2018-11-13
Generating Positive Bounding Boxes for Balanced Training of Object Detectors✓ Link35.655.3Online Fg Bal. Sampling+Hard Negative Mining (ResNet-50)2019-09-21
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices✓ Link34.349.82.15PP-PicoDet-M (416)2021-11-01
M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network✓ Link34.153.715.939.549.3M2Det (ResNet-101, 320x320)2018-11-12
Res2Net: A New Multi-scale Backbone Architecture✓ Link33.753.61438.351.1Faster R-CNN (Res2Net-50)2019-04-02
M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network✓ Link33.252.21538.249.1M2Det (VGG-16, 320x320)2018-11-12
YOLOX: Exceeding YOLO Series in 2021✓ Link32.85.06YOLOX-Tiny (416)2021-07-18
YOLOX: Exceeding YOLO Series in 2021✓ Link25.30.91YOLOX-Nano (416)2021-07-18
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices✓ Link23.50.95NanoDet-M (416, reported by PP-PicoDet)2021-11-01
You Only Learn One Representation: Unified Network for Multiple Tasks✓ Link73.560.640.460.168.7YOLOR-D6 (1280, single-scale, 31 fps)2021-05-10
You Only Learn One Representation: Unified Network for Multiple Tasks✓ Link70.657.437.457.365.2YOLOR-P6 (1280, single-scale, 72 fps)2021-05-10
Focal Modulation Networks✓ Link70.155.8FocalNet-T (SRF, Cascade Mask R-CNN)2022-03-22
Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing✓ Link61.245.624.4R3-CNN (ResNet-50-FPN, GRoIE)2021-04-03
When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism✓ Link42.3Shift-T2022-01-26