ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Benchmark Model Rank Results
2d-human-pose-estimation-on-human-artViTPose-h#3AP: 0.468AP (gt bbox): 0.800
2d-human-pose-estimation-on-human-artViTPose-l#4AP: 0.459AP (gt bbox): 0.789
2d-human-pose-estimation-on-human-artViTpose-b#6AP: 0.410AP (gt bbox): 0.759
2d-human-pose-estimation-on-human-artViTPose-s#8AP: 0.381AP (gt bbox): 0.738
human-pose-keypoint-estimation-on-coco-test-devViTPose (ViTAE-G, ensemble)#1AP: 81.1AP50: 95.0AP75: 88.2AR: 85.6
human-pose-keypoint-estimation-on-coco-test-devViTPose (ViTAE-G)#2AP: 80.9AP50: 94.8AP75: 88.1AR: 85.4
human-pose-keypoint-estimation-on-coco-val2017ViTPose-H (multi-dataset)#9AP: 79.5AR: 84.5
human-pose-keypoint-estimation-on-coco-val2017ViTPose-H#13AP: 79.1AR: 84.1
human-pose-keypoint-estimation-on-coco-val2017ViTPose-L (multi-dataset)#16AP: 78.7AR: 83.8
human-pose-keypoint-estimation-on-coco-val2017ViTPose-L#19AP: 78.3AR: 83.5
human-pose-keypoint-estimation-on-coco-val2017ViTPose-B (Single-task_GT-bbox_256x192)#29AP: 77.3AP50: 93.5AP75: 84.5AR: 80.4
human-pose-keypoint-estimation-on-coco-val2017ViTPose-B (multi-dataset)#32AP: 77.1AR: 82.2
human-pose-keypoint-estimation-on-coco-val2017ViTPose-B (Single-task_Det-bbox_256x192)#46AP: 75.8AP50: 90.7AP75: 83.2AR: 81.1
pose-estimation-on-crowdposeViTPose-G#2AP: 78.3AP50: 85.3AP75: 81.4APM: 86.6AP Hard: 67.9
pose-estimation-on-ochumanViTPose (ViTAE-G, GT bounding boxes)#1Test AP: 93.3Validation AP: 92.8