Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
PIIP-LLaVA (Vicuna-7B, ConvNeXt-L, CLIP-L )
#60
GPT-4 score: 44.7
Params: 7B
Rank counts only results with a code link.