Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models

Benchmark Model Rank Results
robot-manipulation-on-calvinRoboVLMs#3avg. sequence length (D to D): 4.25
robot-manipulation-on-simpler-envRoboVLM#4Visual Matching: 0.563Visual Matching-Pick Coke Can: 0.727
robot-manipulation-on-simplerenv-widow-xRoboVLM#2Average: 0.135Put Spoon on Towel: 0.208Put Carrot on Plate: 0.250