ViTPose++: Vision Transformer for Generic Body Pose Estimation

Benchmark Model Rank Results
2d-human-pose-estimation-on-coco-wholebody-1ViTPose+-H#24WB: 61.2body: 75.9foot: 77.9face: 63.3hand: 54.7
2d-human-pose-estimation-on-coco-wholebody-1ViTPose+-L#26WB: 60.6body: 75.3foot: 77.1face: 63.0hand: 54.2
animal-pose-estimation-on-ap-10kViTPose+-H#1AP: 82.4
animal-pose-estimation-on-ap-10kViTPose+-L#2AP: 80.4
animal-pose-estimation-on-ap-10kViTPose+-B#5AP: 74.5
animal-pose-estimation-on-ap-10kHRNet-w48#6AP: 73.1
animal-pose-estimation-on-ap-10kHRNet-w32#7AP: 72.2
animal-pose-estimation-on-ap-10kViTPose+-S ViT-S#8AP: 71.4
animal-pose-estimation-on-ap-10kSimpleBaseline-ResNet50#9AP: 68.1
human-pose-keypoint-estimation-on-coco-test-devViTPose++-H (multi-dataset)#9AP: 78.5AP50: 93.4AP75: 86.2AR: 83.4
human-pose-keypoint-estimation-on-coco-test-devViTPose-H#12AP: 78.1AP50: 93.3AP75: 85.7AR: 83.1
human-pose-keypoint-estimation-on-coco-val2017ViTPose++-H (multi-dataset)#10AP: 79.4