FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
FlashSloth-HD
#49
GPT-4 score: 49.0
Params: 3.2B
visual-question-answering-on-mm-vet
FlashSloth
#68
GPT-4 score: 41.9
Params: 3.2B
Rank counts only results with a code link.