DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
DeepStack-L-HD (Vicuna-13B)
–
GPT-4 score: 39.3
zero-shot-video-question-answer-on-next-qa
DeepStack-L(7B)
–
Accuracy: 61.0
Rank counts only results with a code link.