Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

Benchmark Model Rank Results
visual-navigation-on-room-to-roomRCM+SIL(no early exploration)–spl: 0.38