Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
LLaVA-1.5+CoS
#88
GPT-4 score: 37.6
Rank counts only results with a code link.