SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Benchmark Model Rank Results
robot-manipulation-on-simpler-envSpatialVLAVisual Matching: 0.719Visual Matching-Pick Coke Can: 0.810
robot-manipulation-on-simplerenv-widow-xSpatialVLAAverage: 0.344Put Spoon on Towel: 0.208Put Carrot on Plate: 0.208