DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Benchmark Model Rank Results
visual-question-answering-on-mm-vetDeepStack-L-HD (Vicuna-13B)GPT-4 score: 39.3
zero-shot-video-question-answer-on-next-qaDeepStack-L(7B)Accuracy: 61.0