| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 81.1 | 95.0 | 88.2 | 85.6 | | | ViTPose (ViTAE-G, ensemble) | 2022-04-26 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 80.9 | 94.8 | 88.1 | 85.4 | | | ViTPose (ViTAE-G) | 2022-04-26 |
| Polarized Self-Attention: Towards High-quality Pixel-wise Regression | ✓ Link | 79.5 | 93.6 | 85.9 | 81.9 | | | UDP-Pose-PSA(384x288) | 2021-07-02 |
| Learning Delicate Local Representations for Multi-Person Pose Estimation | ✓ Link | 79.2 | 94.4 | 87.1 | 84.1 | | | 4xRSN-50 (ensemble) | 2020-03-09 |
| Self-Constrained Inference Optimization on Structural Groups for Human Pose Estimation | | 79.2 | 93.5 | 85.8 | 81.6 | | | SCIO (HRNet-48) | 2022-07-06 |
| Towards High Performance Human Keypoint Detection | ✓ Link | 78.9 | 93.8 | 86.0 | 83.6 | | | CCM+ | 2020-02-03 |
| Polarized Self-Attention: Towards High-quality Pixel-wise Regression | ✓ Link | 78.9 | 93.6 | 85.8 | 81.4 | | | UDP-Pose-PSA(256x192) | 2021-07-02 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 78.7 | | | | | | AID (HRNet-W48plus) | 2020-08-17 |
| Learning Delicate Local Representations for Multi-Person Pose Estimation | ✓ Link | 78.6 | 94.3 | 86.6 | 83.8 | | | 4xRSN-50 | 2020-03-09 |
| PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation | | 78.6 | 93.3 | 86.2 | 83.5 | | | PoseBH-H | 2025-05-23 |
| ViTPose++: Vision Transformer for Generic Body Pose Estimation | ✓ Link | 78.5 | 93.4 | 86.2 | 83.4 | | | ViTPose++-H (multi-dataset) | 2022-12-07 |
| Poseur: Direct Human Pose Regression with Transformers | ✓ Link | 78.3 | 93.5 | 85.9 | | | | Poseur(384x288) | 2022-01-19 |
| Human Pose as Compositional Tokens | ✓ Link | 78.3 | 92.9 | 85.9 | | | | PCT (256x256) | 2023-03-21 |
| ViTPose++: Vision Transformer for Generic Body Pose Estimation | ✓ Link | 78.1 | 93.3 | 85.7 | 83.1 | | | ViTPose-H | 2022-12-07 |
| Poseur: Direct Human Pose Regression with Transformers | ✓ Link | 77.6 | 92.9 | 85.0 | | | | Poseur (HRNet-W48, 384x288) | 2022-01-19 |
| Distribution-Aware Coordinate Representation for Human Pose Estimation | ✓ Link | 77.4 | 92.6 | 84.6 | 82.3 | | | HRNet-W48+DARK | 2019-10-14 |
| Revealing the Dark Secrets of Masked Image Modeling | ✓ Link | 77.2 | | | | | | SwinV2-L 1K-MIM | 2022-05-26 |
| Heatmap Distribution Matching for Human Pose Estimation | | 77.2 | 93.0 | 84.4 | 82.0 | | | HDM (HRNet-W48, 384x288) | 2022-10-03 |
| Heatmap Distribution Matching for Human Pose Estimation | | 77.2 | 93.1 | 84.7 | 82.1 | | | HDM (HRFormer-B, 384x288) | 2022-10-03 |
| Deep High-Resolution Representation Learning for Human Pose Estimation | ✓ Link | 77.0 | 92.7 | 84.5 | 82.0 | | | HRNet-W48 + extra data | 2019-02-25 |
| Lightweight Super-Resolution Head for Human Pose Estimation | ✓ Link | 76.9 | | | 81.8 | | | SRPose (HRFormer-B, 384x288) | 2023-07-31 |
| Graph-PCNN: Two Stage Human Pose Estimation with Graph Pose Refinement | | 76.8 | 92.6 | 84.3 | 81.6 | | | Graph-PCNN (HRNet-W48, 384x288) | 2020-07-21 |
| EvoPose2D: Pushing the Boundaries of 2D Human Pose Estimation using Accelerated Neuroevolution with Weight Transfer | ✓ Link | 76.8 | 92.5 | 84.3 | 81.7 | | | EvoPose2D-L | 2020-11-17 |
| PoseFix: Model-agnostic General Human Pose Refinement Network | ✓ Link | 76.7 | 92.6 | 84.1 | 81.5 | 95.8 | 88.1 | PoseFix + HRNet-W48 | 2018-12-10 |
| Revealing the Dark Secrets of Masked Image Modeling | ✓ Link | 76.7 | | | | | | SwinV2-B 1K-MIM | 2022-05-26 |
| SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation | ✓ Link | 76.7 | 92.8 | 84.4 | 81.6 | | | SHaRPose-Base (384x288) | 2023-12-17 |
| Adaptive Hypergraph Neural Network for Multi-Person Pose Estimation | | 76.6 | 92.4 | 84.3 | 81.5 | | | AD-HNN (HRNet-W48) | 2022-06-28 |
| PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation | | 76.6 | 92.6 | 84.4 | 81.7 | | | PoseBH-B | 2025-05-23 |
| Simple Baselines for Human Pose Estimation and Tracking | ✓ Link | 76.5 | 92.4 | 84.0 | 81.5 | 95.8 | 88.2 | Simple Base+* | 2018-04-17 |
| The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation | ✓ Link | 76.5 | 92.7 | 84.0 | 81.6 | | | HRNet-W48+UDP | 2019-11-18 |
| Inter-image Contrastive Consistency for Multi-Person Pose Estimation | | 76.5 | 92.5 | 83.9 | | | | ICON (HRNet-W48, 384x288) | 2023-06-26 |
| OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation | ✓ Link | 76.4 | 92.6 | 83.7 | 81.2 | | | OmniPose (WASPv2) | 2021-03-18 |
| Adaptive Hypergraph Neural Network for Multi-Person Pose Estimation | | 76.3 | 92.3 | 83.8 | 81.2 | | | AD-HNN (HRNet-W32) | 2022-06-28 |
| Spatial-Aware Regression for Keypoint Localization | ✓ Link | 76.3 | 92.5 | 83.6 | 81.2 | | | SAR (HRNet-W48) | 2024-06-16 |
| Distribution-Aware Coordinate Representation for Human Pose Estimation | ✓ Link | 76.2 | | | | | | DarkPose(384x288) | 2019-10-14 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 76.2 | | | | | | AID (HRNet-W32) | 2020-08-17 |
| HRFormer: High-Resolution Transformer for Dense Prediction | ✓ Link | 76.2 | 92.7 | 83.8 | 81.2 | | | HRFormer-B | 2021-10-18 |
| Rethinking on Multi-Stage Networks for Human Pose Estimation | ✓ Link | 76.1 | 93.4 | 83.8 | 81.6 | 96.3 | 88.1 | MSPN | 2019-01-01 |
| TCFormer: Visual Recognition via Token Clustering Transformer | ✓ Link | 76.1 | 92.4 | 83.7 | | | | RLE + TCFormerV2-Base (384x288) | 2024-07-16 |
| VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation | | 76.1 | 92.9 | 83.9 | | | | VisualCent (ResNet-152) | 2025-04-26 |
| Inter-image Contrastive Consistency for Multi-Person Pose Estimation | | 76.0 | 92.4 | 83.5 | | | | ICON (HRNet-W32, 384x288) | 2023-06-26 |
| Towards Simple and Accurate Human Pose Estimation with Stair Network | | 75.9 | 92.3 | 83.4 | 81.1 | | | STNet 3-stage* (384x288) | 2022-02-18 |
| Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation | ✓ Link | 75.7 | 92.4 | 83.3 | 80.5 | | | MIPNet | 2021-01-27 |
| AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation | ✓ Link | 75.7 | | | | | | AggPose(256x192) | 2022-05-11 |
| Deep Multi-Task Networks For Occluded Pedestrian Pose Estimation | | 75.7 | 90.3 | 76.3 | | | | PPE (ResNeXt-101) | 2022-06-15 |
| DE-HRNet: Detail enhanced high-resolution network for human pose estimation | | 75.7 | 92.3 | 83.2 | 80.9 | | | DE-HRNet-W48 (384x288) | 2025-09-02 |
| DE-HRNet: Detail enhanced high-resolution network for human pose estimation | | 75.6 | 92.4 | 83.3 | 80.7 | | | DE-HRNet-W32 (384x288) | 2025-09-02 |
| Deep High-Resolution Representation Learning for Human Pose Estimation | ✓ Link | 75.5 | 92.5 | 83.3 | 80.5 | | | HRNet-W48 | 2019-02-25 |
| Towards Simple and Accurate Human Pose Estimation with Stair Network | | 75.3 | 92.1 | 82.7 | 80.4 | | | STNet 3-stage (384x288) | 2022-02-18 |
| A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network | | 75.1 | 92.2 | 82.7 | 80.2 | | | Two-stage distillation + PGNN (HRNet-W32) | 2025-08-15 |
| TransPose: Keypoint Localization via Transformer | ✓ Link | 75.0 | 92.2 | 82.3 | | | | TransPose-H-A6 | 2020-12-28 |
| PoseFix: Model-agnostic General Human Pose Refinement Network | ✓ Link | 74.9 | 91.2 | 81.9 | 79.9 | 94.8 | 86.3 | PoseFix (SimpleBaseline ResNet-152) | 2018-12-10 |
| An enhanced real-time human pose estimation method based on modified YOLOv8 framework | | 74.9 | 93.7 | 80.8 | 82.1 | | | CCAM-Person (960) | 2024-04-05 |
| DPIT: Dual-Pipeline Integrated Transformer for Human Pose Estimation | | 74.6 | 91.9 | 82.1 | 79.9 | | | DPIT-L | 2022-09-02 |
| A Context-and-Spatial Aware Network for Multi-Person Pose Estimation | | 74.5 | 91.7 | 82.1 | 80.7 | | | CSANet (ResNet-152, 384x288) | 2019-05-14 |
| Multi-Person Pose Estimation with Enhanced Channel-wise and Spatial Information | | 74.3 | 91.8 | 81.9 | 80.5 | | | CSM+SCARB (ResNet-152, 384x288) | 2019-05-09 |
| Denoising and Selecting Pseudo-Heatmaps for Semi-Supervised Human Pose Estimation | | 74.2 | 92.1 | 82.4 | 79.4 | | | SimpleBaseline-R152 + pseudo-heatmaps (unlabeled COCO + AIC) | 2023-09-29 |
| VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation | | 74.2 | 89.0 | 80.2 | | | | VisualCent (ResNet-101) | 2025-04-26 |
| Lightweight Super-Resolution Head for Human Pose Estimation | ✓ Link | 74.1 | | | 79.1 | | | SRPose (ResNet-50, 384x288) | 2023-07-31 |
| ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search | ✓ Link | 73.9 | 91.7 | 82.0 | 80.4 | | | S-ViPNAS-HRNetW32 | 2021-05-21 |
| Simple Baselines for Human Pose Estimation and Tracking | ✓ Link | 73.7 | 91.9 | 81.1 | 79.0 | | | Flow-based (ResNet-152) | 2018-04-17 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 73.7 | | | | | | AID (ResNet-50) | 2020-08-17 |
| DistilPose: Tokenized Pose Regression with Heatmap Distillation | ✓ Link | 73.7 | 91.6 | 81.1 | | | | DistilPose-L | 2023-03-04 |
| An efficient and accurate 2D human pose estimation method using VTTransPose network | | 73.6 | 91.4 | 81.1 | | | | VTTransPose (256x192) | 2023-07-25 |
| Spatial-Aware Regression for Keypoint Localization | ✓ Link | 73.5 | 91.9 | 80.9 | 78.8 | | | SAR (ResNet-50) | 2024-06-16 |
| RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation | ✓ Link | 73.3 | 91.9 | 80.8 | 77.4 | | | RTMO-l (extra data) | 2023-12-12 |
| Cascaded Pyramid Network for Multi-Person Pose Estimation | ✓ Link | 73.0 | 91.7 | 80.9 | 79.0 | 95.1 | 85.9 | CPN+ [6, 9] | 2017-11-20 |
| Joint Human Pose Estimation and Instance Segmentation with PosePlusSeg | ✓ Link | 72.8 | 88.4 | 78.7 | | | | PosePlusSeg (ResNet-152) | 2022-06-28 |
| Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation | | 72.8 | 92.5 | 81.0 | | | | Group Pose (Swin-L) | 2023-08-14 |
| AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time | ✓ Link | 72.7 | 92.2 | 81.3 | 78.1 | | | FastPose-dcn-hm ResNet-101 (AlphaPose) | 2022-11-07 |
| Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation | ✓ Link | 72.7 | 92.3 | 80.9 | | | | ED-Pose (Swin-L) | 2023-02-03 |
| SDPose: Tokenized Pose Estimation via Circulation-Guide Self-Distillation | ✓ Link | 72.7 | 91.2 | 80.3 | | | | SDPose-S-V2 (self-distillation) | 2024-04-04 |
| A Coarse-Fine Network for Keypoint Localization | | 72.6 | 86.1 | | | | | CFN | 2017-10-22 |
| AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time | ✓ Link | 72.6 | 92.2 | 81.2 | 78.1 | | | FastPose-dcn-hm ResNet-50 (AlphaPose) | 2022-11-07 |
| HEViTPose: High-Efficiency Vision Transformer for Human Pose Estimation | | 72.6 | 92.0 | 80.9 | 78.0 | | | HEViTPose-B | 2023-11-22 |
| RMPE: Regional Multi-person Pose Estimation | ✓ Link | 72.3 | 89.2 | 79.1 | | | | RMPE++ | 2016-12-01 |
| QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query | ✓ Link | 72.3 | 91.5 | 78.7 | | | | QueryPose (HRNet-W48) | 2022-12-15 |
| A Characteristic Function-Based Method for Bottom-Up Human Pose Estimation | | 72.3 | 91.5 | 79.8 | | | | Characteristic Function (HrHRNet-W48, multi-scale) | 2023-06-18 |
| TFPose: Direct Human Pose Estimation with Transformers | | 72.2 | 90.9 | 80.1 | | | | TFPose (ND=6 ResNet-50) | 2021-03-29 |
| QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query | ✓ Link | 72.2 | 92.0 | 78.8 | | | | QueryPose (Swin-L) | 2022-12-15 |
| DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints | ✓ Link | 72.2 | 91.4 | 79.3 | 78.8 | | | DETRPose-X | 2025-06-16 |
| Cascaded Pyramid Network for Multi-Person Pose Estimation | ✓ Link | 72.1 | 91.4 | 80.0 | 78.5 | 95.1 | 85.3 | CPN | 2017-11-20 |
| ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation | | 71.9 | 91.9 | 79.6 | | | | ER-Pose-l (800x800) | 2026-03-09 |
| Learning Quality-aware Representation for Multi-person Pose Regression | | 71.7 | 90.4 | 78.7 | 76.5 | | | CIR&QEM (HRNet-W48, multi-scale) | 2022-01-04 |
| ScaleNAS: One-Shot Learning of Scale-Aware Representations for Visual Recognition | | 71.6 | 90.3 | 78.2 | 76.0 | 92.3 | | HigherHRNet (ScaleNet_P4) | 2020-11-30 |
| RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation | ✓ Link | 71.6 | 91.1 | 79.0 | 75.6 | | | RTMO-l | 2023-12-12 |
| Bottom-Up 2D Pose Estimation via Dual Anatomical Centers for Small-Scale Persons | | 71.5 | 89.1 | 78.5 | | | | Dual Anatomical Centers (HRNet-W48, multi-scale) | 2022-08-25 |
| The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose Estimation | ✓ Link | 71.4 | | | | | | CenterGroup | 2021-10-11 |
| AdaptivePose++: A Powerful Single-Stage Network for Multi-Person Pose Regression | ✓ Link | 71.4 | 90.2 | 78.5 | | | | AdaptivePose++ (HRNet-W48, multi-scale) | 2022-10-08 |
| AdaptivePose: Human Parts as Adaptive Points | ✓ Link | 71.3 | 90.0 | 78.3 | | | | AdaptivePose (HRNet-W48, multi-scale) | 2021-12-27 |
| BAPose: Bottom-Up Pose Estimation with Disentangled Waterfall Representations | ✓ Link | 71.2 | 89.4 | 78.1 | 76.8 | | | BAPose (W48, multi-scale) | 2021-12-20 |
| End-to-End Multi-Person Pose Estimation with Transformers | ✓ Link | 71.2 | 91.4 | 79.6 | | | | PETR (Swin-L, multi-scale) | 2022-06-19 |
| BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation | ✓ Link | 71.2 | 90.8 | 78.6 | 77.1 | | | BoIR (HRNet-W48) | 2023-09-25 |
| DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints | ✓ Link | 71.2 | 91.2 | 78.1 | 78.1 | | | DETRPose-L | 2025-06-16 |
| SIMPLE: SIngle-network with Mimicking and Point Learning for Bottom-up Human Pose Estimation | | 71.1 | 90.2 | 79.4 | | | | SIMPLE-W32 (multi-scale) | 2021-04-06 |
| A Characteristic Function-Based Method for Bottom-Up Human Pose Estimation | | 71.1 | 90.4 | 78.2 | | | | Characteristic Function (HrHRNet-W48) | 2023-06-18 |
| Learning Quality-aware Representation for Multi-person Pose Regression | | 71.0 | 90.2 | 78.2 | 76.0 | | | CIR&QEM (HRNet-W48) | 2022-01-04 |
| Bottom-Up 2D Pose Estimation via Dual Anatomical Centers for Small-Scale Persons | | 71.0 | 89.5 | 78.0 | | | | Dual Anatomical Centers (HRNet-W48) | 2022-08-25 |
| DistilPose: Tokenized Pose Regression with Heatmap Distillation | ✓ Link | 71.0 | 91.0 | 78.9 | | | | DistilPose-S | 2023-03-04 |
| Pose Neural Fabrics Search | ✓ Link | 70.9 | | | | | | PNFS | 2019-09-16 |
| OpenPifPaf: Composite Fields for Semantic Keypoint Detection and Spatio-Temporal Association | ✓ Link | 70.9 | | | | | | OpenPifPaf | 2021-03-03 |
| An improved lightweight high-resolution network based on multi-dimensional weighting for human pose estimation | | 70.9 | 91.0 | 78.3 | 76.8 | | | MDW-HRNet-30 (384x288) | 2023-05-04 |
| FasterPose: A Faster Simple Baseline for Human Pose Estimation | ✓ Link | 70.8 | 91.3 | 78.8 | 76.4 | | | FasterPose (ResNet-50, 256x192) | 2021-07-07 |
| Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation | ✓ Link | 70.8 | 89.7 | 77.8 | | | | LOGO-CAP (HRNet-W48) | 2021-09-08 |
| Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation | ✓ Link | 70.6 | 90.8 | 78.2 | 76.4 | | | Dite-HRNet-30 | 2022-04-22 |
| HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation | ✓ Link | 70.5 | 89.3 | 77.2 | | | | HigherHRNet (HR-Net-48) | 2019-08-27 |
| End-to-End Multi-Person Pose Estimation with Transformers | ✓ Link | 70.5 | 91.5 | 78.7 | | | | PETR (Swin-L) | 2022-06-19 |
| Greit-HRNet: Grouped Lightweight High-Resolution Network for Human Pose Estimation | | 70.5 | 90.6 | 78.1 | 76.1 | | | Greit-HRNet-30 (384x288) | 2024-07-10 |
| LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation | | 70.4 | 91.2 | 77.8 | 75.7 | | | LGM-Pose (384x288) | 2025-06-05 |
| ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search | ✓ Link | 70.3 | 90.7 | 78.8 | 77.3 | | | S-ViPNAS-Res50 | 2021-05-21 |
| Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-Person Human Pose Estimation | ✓ Link | 70.3 | 91.2 | 77.8 | 77.7 | | | KAPAO-L | 2021-11-16 |
| BAPose: Bottom-Up Pose Estimation with Disentangled Waterfall Representations | ✓ Link | 70.3 | 89.6 | 77.5 | 75.4 | | | BAPose (W48) | 2021-12-20 |
| ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation | | 70.3 | 91.4 | 77.9 | | | | ER-Pose-m (800x800) | 2026-03-09 |
| SMPR: Single-Stage Multi-Person Pose Regression | ✓ Link | 70.2 | 89.7 | 77.5 | | | | SMPR (HR-Net-32) | 2020-06-28 |
| Global Relation Modeling and Refinement for Bottom-Up Human Pose Estimation | | 70.2 | 88.9 | 77.2 | 76.1 | 93.1 | 82.2 | GRM (HRNet-W48, multi-scale) | 2023-03-27 |
| Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation | | 70.2 | 90.5 | 77.8 | | | | Group Pose (ResNet-50) | 2023-08-14 |
| X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-Attention | | 70.0 | 90.6 | 77.7 | 75.5 | | | X-HRNet-30 (384x288) | 2023-10-12 |
| Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation | ✓ Link | 69.8 | 90.2 | 77.2 | | | | ED-Pose (ResNet-50) | 2023-02-03 |
| Lite-HRNet: A Lightweight High-Resolution Network | ✓ Link | 69.7 | 90.7 | 77.5 | 75.4 | | | Lite-HRNet-30 | 2021-04-13 |
| MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network | ✓ Link | 69.6 | 86.3 | 76.6 | 73.5 | | | Pose Residual Network | 2018-07-11 |
| Lightweight Human Pose Estimation Using Heatmap-Weighting Loss | ✓ Link | 69.2 | 90.6 | 76.9 | 74.7 | | | Heatmap-Weighting Loss (MobileNetV3, 384x288) | 2022-05-21 |
| Global Relation Modeling and Refinement for Bottom-Up Human Pose Estimation | | 69.1 | 89.0 | 76.0 | 74.8 | 92.7 | 80.7 | GRM (HRNet-W48) | 2023-03-27 |
| DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation | ✓ Link | 69.0 | 89.8 | 76.4 | 74.7 | | | DHRNet (HRNet-W32) | 2024-04-22 |
| Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-Person Human Pose Estimation | ✓ Link | 68.8 | 90.5 | 76.5 | 76.3 | | | KAPAO-M | 2021-11-16 |
| Lightweight Multiperson Pose Estimation With Staggered Alignment Self-Distillation | | 68.8 | | | | | | SASD-L | 2024-04-12 |
| PersonLab: Person Pose Estimation and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model | ✓ Link | 68.7 | 89.0 | 75.4 | | | | PersonLab (multi-scale) | 2018-03-22 |
| Towards Accurate Multi-person Pose Estimation in the Wild | | 68.5 | 87.1 | 75.5 | 73.3 | | | G-RMI* | 2017-01-06 |
| YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss | ✓ Link | 68.5 | 90.3 | 74.8 | 75.0 | | | YOLOv5l6-pose | 2022-04-14 |
| Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation | ✓ Link | 68.4 | 89.9 | 75.8 | 74.4 | | | Dite-HRNet-18 (384x288) | 2022-04-22 |
| BiHRNet: A Binary high-resolution network for Human Pose Estimation | | 68.3 | 90.2 | 76.1 | | | | BiHRNet (384x288) | 2023-11-17 |
| Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose Estimation | ✓ Link | 68.1 | | | 72.1 | 88.2 | | Simple Pose | 2019-11-24 |
| Integral Human Pose Regression | ✓ Link | 67.8 | 88.2 | 74.8 | | | | Integral Pose (ResNet-101, 256x256) | 2017-11-22 |
| 2D Human Pose Estimation with Explicit Anatomical Keypoints Structure Constraints | | 67.7 | 88.3 | 74.6 | | 72.8 | | DEKR + keypoint structure constraints (HRNet-W32) | 2022-12-05 |
| Joint Coordinate Regression and Association For Multi-Person Pose Estimation, A Pure Neural Network Approach | | 67.6 | 90.0 | 75.2 | | | | JCRA (ResNet-50) | 2023-07-03 |
| AdaptivePose: Human Parts as Adaptive Points | ✓ Link | 67.4 | 88.2 | 73.7 | | | | AdaptivePose (DLA-34, multi-scale) | 2021-12-27 |
| PoseTrans: A Simple Yet Effective Pose Transformation Augmentation for Human Pose Estimation | | 67.4 | 88.3 | 73.9 | 72.2 | | | HigherHRNet-W32 + PoseTrans | 2022-08-16 |
| Towards High Performance One-Stage Human Pose Estimation | | 67.1 | 89.0 | 74.0 | | | | One-Stage HPE (ResNet-101) | 2023-01-12 |
| Single-Stage Multi-Person Pose Machines | ✓ Link | 66.9 | 88.5 | 72.9 | | | | SPM | 2019-08-24 |
| Lite-HRNet: A Lightweight High-Resolution Network | ✓ Link | 66.9 | 89.4 | 74.4 | 72.6 | | | Lite-HRNet-18 | 2021-04-13 |
| PifPaf: Composite Fields for Human Pose Estimation | ✓ Link | 66.7 | | | | | | PifPaf (single-scale) | 2019-03-15 |
| YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss | ✓ Link | 66.6 | 89.8 | 73.8 | 73.4 | | | YOLOv5m6-pose | 2022-04-14 |
| PersonLab: Person Pose Estimation and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model | ✓ Link | 66.5 | 88.0 | 72.6 | 71.0 | | | PersonLab (single-scale) | 2018-03-22 |
| Attend to Who You Are: Supervising Self-Attention for Keypoint Detection and Instance-Aware Association | ✓ Link | 66.5 | | | | | | Supervising Self-Attention | 2021-11-25 |
| Towards High Performance One-Stage Human Pose Estimation | | 66.4 | 88.4 | 73.1 | | | | One-Stage HPE (ResNet-50) | 2023-01-12 |
| MovePose: A High-performance Human Pose Estimation Algorithm on Mobile and Edge Devices | | 65.9 | 88.9 | 73.0 | | | | MovePose (flip test) | 2023-08-17 |
| Greedy Offset-Guided Keypoint Grouping for Human Pose Estimation | ✓ Link | 65.6 | | | | | | Hourglass-104 | 2021-07-07 |
| Associative Embedding: End-to-End Learning for Joint Detection and Grouping | ✓ Link | 65.5 | 86.8 | 72.3 | 70.2 | 89.5 | 76.0 | AE | 2016-11-16 |
| MDPose: Real-Time Multi-Person Pose Estimation via Mixture Density Model | | 65.0 | 88.9 | 72.8 | | | | MDPose (ResNet-101) | 2023-02-17 |
| Towards Accurate Multi-person Pose Estimation in the Wild | | 64.9 | 85.5 | 71.3 | 69.7 | 88.7 | 75.5 | G-RMI | 2017-01-06 |
| DirectPose: Direct End-to-End Multi-Person Pose Estimation | ✓ Link | 64.8 | 87.8 | 71.1 | | | | DirectPose (ResNet-101, multi-scale) | 2019-11-18 |
| Revisiting Unreasonable Effectiveness of Data in Deep Learning Era | ✓ Link | 64.4 | 85.7 | 70.7 | | | | Faster R-CNN (ImageNet+300M) | 2017-07-10 |
| OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields | ✓ Link | 64.2 | 86.2 | 70.1 | | | | OpenPose | 2018-12-18 |
| Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-Person Human Pose Estimation | ✓ Link | 63.8 | 88.4 | 70.4 | 71.2 | | | KAPAO-S | 2021-11-16 |
| DirectPose: Direct End-to-End Multi-Person Pose Estimation | ✓ Link | 63.3 | 86.7 | 69.4 | | | | DirectPose (ResNet-101) | 2019-11-18 |
| Mask R-CNN | ✓ Link | 63.1 | 87.3 | 68.7 | | | | Mask-RCNN | 2017-03-20 |
| Associative Embedding: End-to-End Learning for Joint Detection and Grouping | ✓ Link | 62.8 | | | | | | AE (single-scale) | 2016-11-16 |
| Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields | ✓ Link | 61.8 | 84.9 | 67.5 | 66.5 | 87.2 | | CMU-Pose | 2016-11-24 |
| RMPE: Regional Multi-person Pose Estimation | ✓ Link | 61.8 | 83.7 | 69.8 | | | | RMPE | 2016-12-01 |
| Lite Pose: Efficient Architecture Design for 2D Human Pose Estimation | ✓ Link | 56.7 | | | | | | LitePose-S | 2022-05-03 |