| Sapiens: Foundation for Human Vision Models | ✓ Link | 82.2 | | | | Sapiens-2B | 2024-08-22 |
| Sapiens: Foundation for Human Vision Models | ✓ Link | 82.1 | | | | Sapiens-1B | 2024-08-22 |
| Sapiens: Foundation for Human Vision Models | ✓ Link | 81.2 | | | | Sapiens-0.6B | 2024-08-22 |
| SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation | ✓ Link | 81.2 | | | 85.3 | SDPose (diffusion, 1024x768) | 2025-09-29 |
| DAREPose: Causally Inspired Risk-Adaptive Correction for Robust Whole-Body Pose Estimation | | 80.4 | | | 82.8 | DAREPose-l | 2026-08-26 |
| RSPose: Ranking Based Losses for Human Pose Estimation | | 79.9 | 92.0 | 86.4 | 84.7 | RSPose (ViTPose-H) | 2025-11-17 |
| Poseur: Direct Human Pose Regression with Transformers | ✓ Link | 79.6 | | | | Poseur(384x288) | 2022-01-19 |
| Sapiens: Foundation for Human Vision Models | ✓ Link | 79.6 | | | | Sapiens-0.3B | 2024-08-22 |
| OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation | ✓ Link | 79.5 | 93.6 | 85.9 | 81.9 | OmniPose (WASPv2) | 2021-03-18 |
| Polarized Self-Attention: Towards High-quality Pixel-wise Regression | ✓ Link | 79.5 | | | | UDP-Pose-PSA(384x288) | 2021-07-02 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 79.5 | | | 84.5 | ViTPose-H (multi-dataset) | 2022-04-26 |
| Self-Constrained Inference Optimization on Structural Groups for Human Pose Estimation | | 79.5 | 93.7 | 86.0 | 81.6 | SCIO (HRNet-48) | 2022-07-06 |
| PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation | | 79.5 | 91.9 | 85.8 | 84.5 | PoseBH-H | 2025-05-23 |
| ViTPose++: Vision Transformer for Generic Body Pose Estimation | ✓ Link | 79.4 | | | | ViTPose++-H (multi-dataset) | 2022-12-07 |
| DAREPose: Causally Inspired Risk-Adaptive Correction for Robust Whole-Body Pose Estimation | | 79.3 | | | 81.7 | DAREPose-m | 2026-08-26 |
| From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models | ✓ Link | 79.2 | 92.7 | | | ReChannel-9B | 2026-07-07 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 79.1 | | | | AID (HRNet-W48plus) | 2020-08-17 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 79.1 | | | 84.1 | ViTPose-H | 2022-04-26 |
| Harnessing Diffusion Models for Visual Perception with Meta Prompts | ✓ Link | 79.0 | | | | MetaPrompt-SD | 2023-12-22 |
| Heatmap Distribution Matching for Human Pose Estimation | | 78.9 | 92.6 | 85.4 | 83.3 | HDM (HRFormer-B, 384x288) | 2022-10-03 |
| Poseur: Direct Human Pose Regression with Transformers | ✓ Link | 78.8 | 91.6 | 85.1 | | Poseur (HRNet-W48, 384x288) | 2022-01-19 |
| Heatmap Distribution Matching for Human Pose Estimation | | 78.8 | 92.5 | 85.1 | 83.1 | HDM (HRNet-W48, 384x288) | 2022-10-03 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 78.7 | | | 83.8 | ViTPose-L (multi-dataset) | 2022-04-26 |
| Hulk: A Universal Knowledge Translator for Human-Centric Tasks | ✓ Link | 78.7 | | | | Hulk(Finetune, ViT-L) | 2023-12-04 |
| CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation | ✓ Link | 78.5 | | | 81.1 | CIGPose-l (384x288) | 2026-03-10 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 78.3 | | | 83.5 | ViTPose-L | 2022-04-26 |
| On the Calibration of Human Pose Estimation | | 78.1 | 93.7 | 85.0 | 80.4 | CCNet (ViTPose-B_GT-bbox_256x192) | 2023-11-28 |
| From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models | ✓ Link | 78.0 | 91.4 | | | ReChannel-4B | 2026-07-07 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 77.8 | | | | AID (HRNet-W32) | 2020-08-17 |
| DPIT: Dual-Pipeline Integrated Transformer for Human Pose Estimation | | 77.8 | 93.6 | 84.8 | 80.3 | DPIT-L (GT boxes) | 2022-09-02 |
| Rethinking pose estimation in crowds: overcoming the detection information-bottleneck and ambiguity | ✓ Link | 77.8 | | | | BUCTD (PETR, with generative sampling) | 2023-06-13 |
| PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment | ✓ Link | 77.8 | 94.4 | 85.4 | 80.8 | PoseLLM (GT boxes) | 2025-07-12 |
| Learning Structure-Guided Diffusion Model for 2D Human Pose Estimation | | 77.6 | | | 82.4 | DiffusionPose (HRNet-W48, 384x288) | 2023-06-29 |
| CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation | ✓ Link | 77.6 | | | 80.3 | CIGPose-l (256x192) | 2026-03-10 |
| EvoPose2D: Pushing the Boundaries of 2D Human Pose Estimation using Accelerated Neuroevolution with Weight Transfer | ✓ Link | 77.5 | | | | EvoPose2D-L(512x384) | 2020-11-17 |
| Hulk: A Universal Knowledge Translator for Human-Centric Tasks | ✓ Link | 77.5 | | | | Hulk(Finetune, ViT-B) | 2023-12-04 |
| SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation | ✓ Link | 77.4 | 91.0 | 84.1 | 82.4 | SHaRPose-Base (384x288) | 2023-12-17 |
| PoseFix: Model-agnostic General Human Pose Refinement Network | ✓ Link | 77.3 | 90.9 | 83.5 | 82.0 | PoseFix + HRNet-W48 | 2018-12-10 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 77.3 | 93.5 | 84.5 | 80.4 | ViTPose-B (Single-task_GT-bbox_256x192) | 2022-04-26 |
| I^2R-Net: Intra- and Inter-Human Relation Network for Multi-Person Pose Estimation | ✓ Link | 77.3 | 91.0 | 83.6 | 82.1 | I²R-Net (1st stage:HRFormer-B) | 2022-06-22 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 77.3 | 91.4 | 84.0 | 82.2 | MogaNet-B (384x288) | 2022-11-07 |
| MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection | | 77.3 | 90.8 | 83.4 | 82.1 | MamKPD-L | 2024-12-02 |
| PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation | | 77.3 | 90.8 | 84.2 | 82.4 | PoseBH-B | 2025-05-23 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 77.1 | | | 82.2 | ViTPose-B (multi-dataset) | 2022-04-26 |
| HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining | ✓ Link | 77.1 | | | | PATH (Partial FT) | 2023-03-10 |
| Learning Structure-Guided Diffusion Model for 2D Human Pose Estimation | | 77.1 | | | 81.8 | DiffusionPose (HRNet-W32, 384x288) | 2023-06-29 |
| TCFormer: Visual Recognition via Token Clustering Transformer | ✓ Link | 77.1 | 91.0 | 83.7 | 81.5 | RLE + TCFormerV2-Base (384x288) | 2024-07-16 |
| Waterfall Transformer for Multi-person Pose Estimation | | 77.1 | 91.1 | 84.1 | 82.0 | WTPose (Swin-B, 384x288) | 2024-11-28 |
| Towards Simple and Accurate Human Pose Estimation with Stair Network | | 76.8 | 90.7 | 83.3 | 81.8 | STNet 3-stage* (384x288) | 2022-02-18 |
| X-Pose: Detecting Any Keypoints | ✓ Link | 76.8 | | | | X-Pose-T (Swin-L) | 2023-10-12 |
| BTranspose: Bottleneck Transformers for Human Pose Estimation with Self-Supervised Pre-Training | | 76.7 | | | 79.3 | BTranspose C3A1(4)-Dino-Large | 2022-04-21 |
| Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks | ✓ Link | 76.6 | | | 81.5 | SOLIDER (swin-B) | 2023-03-30 |
| ProbPose: A Probabilistic Approach to 2D Human Pose Estimation | | 76.6 | | | | ProbPose-s (GT boxes) | 2024-12-03 |
| DE-HRNet: Detail enhanced high-resolution network for human pose estimation | | 76.6 | 90.4 | 83.3 | 81.6 | DE-HRNet-W48 (384x288) | 2025-09-02 |
| RSPose: Ranking Based Losses for Human Pose Estimation | | 76.6 | 90.8 | 83.4 | 81.5 | RSPose (SimCC HRNet-W48) | 2025-11-17 |
| CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation | ✓ Link | 76.6 | | | 79.3 | CIGPose-m | 2026-03-10 |
| Lightweight Human Pose Estimation Algorithm Based on Polarized Self-Attention | | 76.4 | 92.6 | 83.6 | 81.1 | LPNet-W48 (384x288) | 2022-04-29 |
| AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation | ✓ Link | 76.4 | | | | AggPose(256x192) | 2022-05-11 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 76.4 | 91.0 | 83.3 | 81.4 | MogaNet-S (384x288) | 2022-11-07 |
| DE-HRNet: Detail enhanced high-resolution network for human pose estimation | | 76.4 | 90.6 | 83.2 | 81.5 | DE-HRNet-W32 (384x288) | 2025-09-02 |
| Deep High-Resolution Representation Learning for Human Pose Estimation | ✓ Link | 76.3 | | | | HRNet-48(384x288) | 2019-02-25 |
| Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation | ✓ Link | 76.3 | | | | MIPNet(384x288) | 2021-01-27 |
| RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose | ✓ Link | 76.3 | | | | RTMPose-l (+AIC) | 2023-03-13 |
| Towards Simple and Accurate Human Pose Estimation with Stair Network | | 76.2 | 90.4 | 82.4 | 81.2 | STNet 3-stage (384x288) | 2022-02-18 |
| A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network | | 76.2 | 90.8 | 82.8 | 81.1 | Two-stage distillation + PGNN (HRNet-W32) | 2025-08-15 |
| MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection | | 76.1 | 90.6 | 82.6 | 80.9 | MamKPD-B | 2024-12-02 |
| Lightweight Super-Resolution Head for Human Pose Estimation | ✓ Link | 75.9 | | | 81.0 | SRPose (HRNet-W32, 256x192) | 2023-07-31 |
| EfficientPose: A Lightweight And Efficient Model with Transformer For Human Pose Estimation | | 75.9 | 90.5 | 82.3 | 81.1 | EfficientPose-W48 (256x192) | 2023-11-03 |
| EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation | | 75.9 | 92.4 | 82.7 | 81.2 | ECPose-X (Objects365 pretrain) | 2026-03-19 |
| Deep High-Resolution Representation Learning for Human Pose Estimation | ✓ Link | 75.8 | | | | HRNet-W32 (384x288) | 2019-02-25 |
| TransPose: Keypoint Localization via Transformer | ✓ Link | 75.8 | | | | TransPose(256x192) | 2020-12-28 |
| Removing the Bias of Integral Pose Regression | | 75.8 | | | | Bias (HRNet_256x192) | 2021-01-01 |
| ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation | ✓ Link | 75.8 | 90.7 | 83.2 | 81.1 | ViTPose-B (Single-task_Det-bbox_256x192) | 2022-04-26 |
| Lightweight Human Pose Estimation Algorithm Based on Polarized Self-Attention | | 75.8 | 93.2 | 84.5 | 80.5 | LPNet-W32 (384x288) | 2022-04-29 |
| Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation | ✓ Link | 75.8 | 92.3 | 82.9 | | ED-Pose (Swin-L, Objects365 pretrain) | 2023-02-03 |
| FasterPose: A Faster Simple Baseline for Human Pose Estimation | ✓ Link | 75.6 | | | | FasterPose (ResNet-152, 384x288) | 2021-07-07 |
| Lightweight Super-Resolution Head for Human Pose Estimation | ✓ Link | 75.6 | | | 80.7 | SRPose (HRFormer-S, 256x192) | 2023-07-31 |
| SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation | ✓ Link | 75.5 | 90.6 | 82.3 | 80.8 | SHaRPose-Base (256x192) | 2023-12-17 |
| Poseur: Direct Human Pose Regression with Transformers | ✓ Link | 75.4 | 90.5 | 82.2 | | Poseur (ResNet-50) | 2022-01-19 |
| Deep High-Resolution Representation Learning for Human Pose Estimation | ✓ Link | 75.3 | | | | HRNet (256x192) | 2019-02-25 |
| AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation | ✓ Link | 75.3 | | | | AID (ResNet-50) | 2020-08-17 |
| PPT: token-Pruned Pose Transformer for monocular and multi-view human pose estimation | ✓ Link | 75.2 | 89.8 | 81.7 | 80.4 | PPT-L/D6 (HRNet-W48, 256x192) | 2022-09-16 |
| FasterPose: A Faster Simple Baseline for Human Pose Estimation | ✓ Link | 75.0 | | | | FasterPose (ResNet-101, 384x288) | 2021-07-07 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 74.9 | | | 80.1 | MogaNet-S (256x192) | 2022-11-07 |
| GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation | ✓ Link | 74.9 | | | 80.0 | GTPT-B | 2024-07-15 |
| AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation | | 74.9 | 90.5 | 81.9 | 79.8 | AgentPose-M | 2025-01-14 |
| AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation | | 74.9 | 90.1 | 81.8 | 79.9 | DWPose-M (COCO body) | 2025-01-14 |
| RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose | ✓ Link | 74.8 | | | | RTMPose-l (COCO only) | 2023-03-13 |
| Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation | | 74.8 | 91.6 | 82.1 | | Group Pose (Swin-L) | 2023-08-14 |
| EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation | | 74.8 | 92.2 | 81.5 | 80.1 | ECPose-X | 2026-03-19 |
| An efficient and accurate 2D human pose estimation method using VTTransPose network | | 74.6 | | | 78.5 | VTTransPose (256x192) | 2023-07-25 |
| EfficientPose: A Lightweight And Efficient Model with Transformer For Human Pose Estimation | | 74.5 | 89.6 | 81.2 | 79.5 | EfficientPose-W32 (256x192) | 2023-11-03 |
| Joint Human Pose Estimation and Instance Segmentation with PosePlusSeg | ✓ Link | 74.4 | 89.4 | 74.8 | | PosePlusSeg (ResNet-152) | 2022-06-28 |
| PPT: token-Pruned Pose Transformer for monocular and multi-view human pose estimation | ✓ Link | 74.4 | 89.6 | 80.9 | 79.6 | PPT-B (HRNet-W32, 256x192) | 2022-09-16 |
| DistilPose: Tokenized Pose Regression with Heatmap Distillation | ✓ Link | 74.4 | | | | DistilPose-L | 2023-03-04 |
| Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation | ✓ Link | 74.3 | 91.5 | 81.6 | | ED-Pose (Swin-L) | 2023-02-03 |
| A Characteristic Function-Based Method for Bottom-Up Human Pose Estimation | | 73.7 | 89.9 | 79.6 | | Characteristic Function (HrHRNet-W48, multi-scale) | 2023-06-18 |
| SDPose: Tokenized Pose Estimation via Circulation-Guide Self-Distillation | ✓ Link | 73.7 | 89.6 | 80.4 | 79.1 | SDPose-B (self-distillation) | 2024-04-04 |
| QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query | ✓ Link | 73.6 | 90.3 | 79.7 | | QueryPose (HRNet-W48) | 2022-12-15 |
| GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation | ✓ Link | 73.6 | | | 78.9 | GTPT-S | 2024-07-15 |
| EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation | | 73.5 | 91.7 | 79.9 | 78.8 | ECPose-L | 2026-03-19 |
| QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query | ✓ Link | 73.3 | 91.3 | 79.5 | | QueryPose (Swin-L) | 2022-12-15 |
| DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints | ✓ Link | 73.3 | 90.5 | 79.4 | 79.4 | DETRPose-X | 2025-06-16 |
| MogaNet: Multi-order Gated Aggregation Network | ✓ Link | 73.2 | 90.1 | 81.0 | 78.8 | MogaNet-T (256x192) | 2022-11-07 |
| Empowering Efficient Human Pose Estimation with Semantic Splitting | | 73.1 | | | | SplitBase ResNet-50 (384x288) | 2023-06-14 |
| Bottom-Up 2D Pose Estimation via Dual Anatomical Centers for Small-Scale Persons | | 73.0 | 88.3 | 79.1 | | Dual Anatomical Centers (HRNet-W48, multi-scale) | 2022-08-25 |
| Neural Interactive Keypoint Detection | ✓ Link | 73.0 | 90.4 | 80.0 | | Click-Pose (ResNet-50, no clicks) | 2023-08-20 |
| An improved lightweight high-resolution network based on multi-dimensional weighting for human pose estimation | | 72.9 | 91.6 | 80.4 | 77.8 | MDW-HRNet-30 (384x288) | 2023-05-04 |
| BAPose: Bottom-Up Pose Estimation with Disentangled Waterfall Representations | ✓ Link | 72.7 | 88.6 | 79.1 | 77.9 | BAPose (W48, multi-scale) | 2021-12-20 |
| DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints | ✓ Link | 72.7 | 91.0 | 79.2 | 78.7 | DETRPose-L | 2025-06-16 |
| PE-former: Pose Estimation Transformer | ✓ Link | 72.6 | | | 79.4 | PEFORMER-Xcit-dino-p8 | 2021-12-09 |
| A Characteristic Function-Based Method for Bottom-Up Human Pose Estimation | | 72.5 | 89.3 | 79.1 | | Characteristic Function (HrHRNet-W48) | 2023-06-18 |
| BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation | ✓ Link | 72.5 | 89.9 | 79.1 | 78.3 | BoIR (HRNet-W48) | 2023-09-25 |
| Spatial-Aware Regression for Keypoint Localization | ✓ Link | 72.5 | 92.3 | 79.9 | 76.1 | SAR (ResNet-50, 256x192) | 2024-06-16 |
| Learning Quality-aware Representation for Multi-person Pose Regression | | 72.4 | 89.1 | 79.0 | 76.4 | CIR&QEM (HRNet-W48) | 2022-01-04 |
| Simple Baselines for Human Pose Estimation and Tracking | ✓ Link | 72.2 | | | | SimpleBaseline (ResNet-50, 384x288) | 2018-04-17 |
| Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation | ✓ Link | 72.2 | 88.9 | 78.9 | | LOGO-CAP (HRNet-W48) | 2021-09-08 |
| Bottom-Up 2D Pose Estimation via Dual Anatomical Centers for Small-Scale Persons | | 72.1 | 88.3 | 78.2 | | Dual Anatomical Centers (HRNet-W48) | 2022-08-25 |
| Lightweight Human Pose Estimation Based on Self-Attention Mechanism | | 72.1 | 89.6 | 79.5 | | WGNet | 2023-03-21 |
| AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation | | 72.1 | 89.6 | 79.3 | 77.2 | AgentPose-S | 2025-01-14 |
| LAPX: Lightweight Hourglass Network with Global Context | ✓ Link | 72.1 | | | | LAPX (384x288) | 2025-12-18 |
| Attention-Enhanced Lightweight Hourglass Network for Human Pose Estimation | | 72.07 | | | | LAP [x2] | 2024-12-09 |
| Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation | | 72.0 | 89.4 | 79.1 | | Group Pose (ResNet-50) | 2023-08-14 |
| Ultralytics YOLOv8 | ✓ Link | 71.6 | 91.2 | | | YOLOv8x-pose-p6 | 2023-01-10 |
| Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation | ✓ Link | 71.6 | 89.6 | 78.1 | | ED-Pose (ResNet-50) | 2023-02-03 |
| DistilPose: Tokenized Pose Regression with Heatmap Distillation | ✓ Link | 71.6 | | | | DistilPose-S | 2023-03-04 |
| Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models | ✓ Link | 71.6 | 91.6 | | | YOLO26x-pose | 2026-06-02 |
| BAPose: Bottom-Up Pose Estimation with Disentangled Waterfall Representations | ✓ Link | 71.5 | 88.7 | 77.8 | 76.2 | BAPose (W48) | 2021-12-20 |
| Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation | ✓ Link | 71.5 | 88.9 | 78.2 | 77.2 | Dite-HRNet-30 (384x288) | 2022-04-22 |
| Greit-HRNet: Grouped Lightweight High-Resolution Network for Human Pose Estimation | | 71.4 | 88.8 | 78.0 | 77.3 | Greit-HRNet-30 (384x288) | 2024-07-10 |
| Human Pose Regression with Residual Log-likelihood Estimation | ✓ Link | 71.3 | 88.9 | 78.3 | | RLE (256x192) | 2021-07-23 |
| Global Relation Modeling and Refinement for Bottom-Up Human Pose Estimation | | 71.2 | 87.5 | 77.7 | 77.0 | GRM (HRNet-W48, multi-scale) | 2023-03-27 |
| OpenPifPaf: Composite Fields for Semantic Keypoint Detection and Spatio-Temporal Association | ✓ Link | 71.0 | | | | OpenPifPaf | 2021-03-03 |
| LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation | | 71.0 | 90.5 | 78.1 | 76.5 | LGM-Pose (384x288) | 2025-06-05 |
| AdaptivePose++: A Powerful Single-Stage Network for Multi-Person Pose Regression | ✓ Link | 70.8 | 88.3 | 77.0 | | AdaptivePose++ (HRNet-W48, 800) | 2022-10-08 |
| BiHRNet: A Binary high-resolution network for Human Pose Estimation | | 70.8 | 91.5 | 78.3 | | BiHRNet (384x288) | 2023-11-17 |
| X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-Attention | | 70.6 | 88.9 | 77.7 | 76.1 | X-HRNet-30 (384x288) | 2023-10-12 |
| Simple Baselines for Human Pose Estimation and Tracking | ✓ Link | 70.4 | | | | SimpleBaseLine (256x192) | 2018-04-17 |
| Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models | ✓ Link | 70.4 | 90.5 | | | YOLO26l-pose | 2026-06-02 |
| Global Relation Modeling and Refinement for Bottom-Up Human Pose Estimation | | 70.1 | 88.0 | 76.5 | 75.5 | GRM (HRNet-W48) | 2023-03-27 |
| Lightweight Multiperson Pose Estimation With Staggered Alignment Self-Distillation | | 70.1 | | | | SASD-L | 2024-04-12 |
| Lightweight Human Pose Estimation Using Heatmap-Weighting Loss | ✓ Link | 69.9 | 88.8 | 77.5 | 75.5 | Heatmap-Weighting Loss (MobileNetV3, 384x288) | 2022-05-21 |
| 2D Human Pose Estimation with Explicit Anatomical Keypoints Structure Constraints | | 69.8 | 88.2 | 76.8 | | PointSetAnchor + keypoint structure constraints (HRNet-W32) | 2022-12-05 |
| Lightweight Joint Loss 2D Pose Estimation Network Based on CM-RTMPose | | 69.8 | 90.6 | 77.5 | 71.6 | CM-RTMPose-t | 2023-07-26 |
| Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models | ✓ Link | 69.5 | 91.1 | | | YOLO11x-pose | 2026-06-02 |
| YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss | ✓ Link | 69.4 | 90.2 | 76.1 | 75.9 | YOLOv5l6-pose | 2022-04-14 |
| Ultralytics YOLOv8 | ✓ Link | 69.2 | 90.2 | | | YOLOv8x-pose | 2023-01-10 |
| Joint Coordinate Regression and Association For Multi-Person Pose Estimation, A Pure Neural Network Approach | | 69.2 | 89.4 | 76.7 | | JCRA (ResNet-50) | 2023-07-03 |
| Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models | ✓ Link | 68.8 | 89.9 | | | YOLO26m-pose | 2026-06-02 |
| PoseTrans: A Simple Yet Effective Pose Transformation Augmentation for Human Pose Estimation | | 68.4 | 87.1 | 74.8 | 72.9 | HigherHRNet-W32 + PoseTrans | 2022-08-16 |
| Towards High Performance One-Stage Human Pose Estimation | | 68.3 | 88.0 | 74.8 | | One-Stage HPE (ResNet-101) | 2023-01-12 |
| Towards High Performance One-Stage Human Pose Estimation | | 68.1 | 88.0 | 74.5 | | One-Stage HPE (ResNet-50) | 2023-01-12 |
| Ultralytics YOLOv8 | ✓ Link | 67.6 | 90.0 | | | YOLOv8l-pose | 2023-01-10 |
| YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss | ✓ Link | 67.4 | 89.1 | 73.7 | 73.9 | YOLOv5m6-pose | 2022-04-14 |
| DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints | ✓ Link | 67.0 | 88.1 | 72.9 | 73.5 | DETRPose-S | 2025-06-16 |
| Non-local Neural Networks | ✓ Link | 66.5 | | | | Mask R-CNN + NL blocks (4 in head, 1 in backbone) | 2017-11-21 |
| Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models | ✓ Link | 66.1 | 89.9 | | | YOLO11l-pose | 2026-06-02 |
| MDPose: Real-Time Multi-Person Pose Estimation via Mixture Density Model | | 65.2 | | | | MDPose (ResNet-101) | 2023-02-17 |
| Ultralytics YOLOv8 | ✓ Link | 65.0 | 88.8 | | | YOLOv8m-pose | 2023-01-10 |
| InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose Estimation | ✓ Link | 63.1 | | | | InsPose | 2021-07-19 |
| ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation | | 60.9 | 86.8 | 68.8 | 65.9 | ER-Pose-n (960x960) | 2026-03-09 |
| Ultralytics YOLOv8 | ✓ Link | 60.0 | 86.2 | | | YOLOv8s-pose | 2023-01-10 |
| DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks | | 57.8 | | | | DICEPTION | 2025-02-24 |
| Lite Pose: Efficient Architecture Design for 2D Human Pose Estimation | ✓ Link | 56.8 | | | | LitePose-S | 2022-05-03 |
| DIR-BHRNet: A Lightweight Network for Real-time Vision-based Multi-person Pose Estimation on Smartphones | ✓ Link | 50.5 | | | | DIR-BHRNet-32 | 2024-07-01 |
| Ultralytics YOLOv8 | ✓ Link | 50.4 | 80.1 | | | YOLOv8n-pose | 2023-01-10 |