ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding

Benchmark Model Rank Results
object-detection-on-coco-val2017ParFormer-M (Mask R-CNN 1x)–box AP: 40.7AP50: 63.3AP75: 44.2Params (M): 42.2
object-detection-on-coco-val2017ParFormer-T (Mask R-CNN 1x)–box AP: 37.3AP50: 59.0AP75: 40.4Params (M): 25.7