Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Benchmark Model Rank Results
meme-classification-on-hateful-memesPaLI-X-VPDROC-AUC: 0.892
object-counting-on-tallyqa-complexPaLI-X-VPDAccuracy: 76.6
object-counting-on-tallyqa-simplePaLI-X-VPDAccuracy: 86.2
visual-question-answering-on-a-okvqaPaLI-X-VPDMC Accuracy: 80.4DA VQA Score: 68.2
visual-question-answering-on-gqa-test-devPaLI-X-VPDAccuracy: 67.3
visual-question-answering-on-ok-vqaPaLI-X-VPDAccuracy: 66.8