All in Tokens: Unifying Output Space of Visual Tasks via Soft Token

Benchmark Model Rank Results
monocular-depth-estimation-on-nyu-depth-v2AiT-P(SwinV2-L)#22absolute relative error: 0.076RMSE: 0.275log 10: 0.033