TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action

Benchmark Model Rank Results
visual-question-answering-on-mm-vetTACO (Qwen2-7B / SigLIP)#46GPT-4 score: 50.9
visual-question-answering-on-mm-vetTACO (LLaMA3-8B / SigLIP)#57GPT-4 score: 45.7
visual-question-answering-on-mm-vetTACO (LLaMA3-8B / CLIP)#58GPT-4 score: 45.2