Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
Mono-InternVL-2B
–
GPT-4 score: 40.1
Params: 2B
Rank counts only results with a code link.