Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Open paper
Benchmark
Model
Rank
Results
3d-question-answering-3d-qa-on-scanqa-test-w
Video-3D LLM
#2
Exact Match: 30.1
CIDEr: 102.1
3d-question-answering-3d-qa-on-sqa3d
Video-3D LLM
#1
Exact Match: 58.6
Rank counts only results with a code link.