Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Benchmark Model Rank Results
visual-question-answering-on-mm-vetGPT-4o +text rationale +IoTGPT-4 score: 72.2