ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

Benchmark Model Rank Results
chat-based-image-retrieval-on-visdialImageScope (CLIP-ViT-L/14)#1Hits@10 on 10 Round: 79.89
zero-shot-composed-image-retrieval-zs-cir-on-circoImageScope (CLIP-ViT-L/14)#12mAP@10: 29.23MAP@5: 28.36mAP@50: 31.88mAP@25: 30.81
zero-shot-composed-image-retrieval-zs-cir-on-cirrImageScope (CLIP-ViT-L/14)#3R@1: 39.37R@5: 67.54R@10: 78.05R@50: 92.94
zero-shot-composed-image-retrieval-zs-cir-on-fashion-iqImageScope (CLIP-ViT-L/14)#40R@10: 31.36R@50: 50.78