ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Benchmark Model Rank Results
visual-question-answering-on-v-benchLLaVA-OneVision7B w. ZoomEye#1Accuracy: 90.58