LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Benchmark Model Rank Results
video-based-generative-performanceLLaMA Adapter#22mean: 2.16Correctness of Information: 2.03Detail Orientation: 2.32
video-based-generative-performance-1LLaMA Adapter#17gpt-score: 2.03
video-based-generative-performance-2LLaMA Adapter#17gpt-score: 2.15
video-based-generative-performance-3LLaMA Adapter#17gpt-score: 2.30
video-based-generative-performance-4LLaMA Adapter#17gpt-score: 2.32
video-based-generative-performance-5LLaMA Adapter#15gpt-score: 1.98
video-question-answering-on-activitynet-qaLLaMA Adapter V2#26Accuracy: 34.2Confidence score: 2.7
visual-question-answering-on-mm-vetLLaMA-Adapter v2-7B#140GPT-4 score: 31.4±0.1Params: 7B
visual-question-answering-vqa-on-core-mmLLaMA-Adapter V2#6Overall score: 30.46Deductive: 28.7Abductive: 46.12
zeroshot-video-question-answer-on-activitynetLLaMA Adapter#25Accuracy: 34.2Confidence Score: 2.7
zeroshot-video-question-answer-on-msrvtt-qaLLaMA Adapter-7B#28Accuracy: 43.8Confidence Score: 2.7
zeroshot-video-question-answer-on-msvd-qaLLaMA Adapter-7B#26Accuracy: 54.9Confidence Score: 3.1