Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Benchmark Model Rank Results
object-detection-on-coco-minivalPVT-Large (RetinaNet 3x,MS)#140box AP: 43.4AP50: 63.6AP75: 46.1APS: 26.1APM: 46.0APL: 59.5
object-detection-on-coco-minivalPVT-Large (RetinaNet 1x)#151box AP: 42.6AP50: 63.7AP75: 45.4APS: 25.8APM: 46.0APL: 58.4
semantic-segmentation-on-densepassPVT (Tiny, FPN)#26mIoU: 31.20%
semantic-segmentation-on-synpassPVT#4mIoU: 32.68%