Text-image Alignment for Diffusion-based Perception

Benchmark Model Rank Results
monocular-depth-estimation-on-nyu-depth-v2TADP#16absolute relative error: 0.062RMSE: 0.225log 10: 0.027
semantic-segmentation-on-ade20kTADP#42Validation mIoU: 55.9
semantic-segmentation-on-nighttime-drivingTADP#1mIoU: 60.8
semantic-segmentation-on-pascal-voc-2012-valTADP#2mIoU: 87.11%
weakly-supervised-object-detection-on-1TADP#1MAP: 72.2
weakly-supervised-object-detection-on-comic2kTADP#2MAP: 57.4