ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-v-bench
LLaVA-OneVision7B w. ZoomEye
#1
Accuracy: 90.58
Rank counts only results with a code link.