Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

Benchmark Model Rank Results
visual-question-answering-on-mm-vetLLaVA1.5-13B-MDAGPT-4 score: 39.90Params: 13B
visual-question-answering-on-mm-vetLLaVA1.5-7B-MDAGPT-4 score: 35.20Params: 7B