Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning

Benchmark Model Rank Results
visual-entailment-on-snli-ve-testSOHO#5Accuracy: 84.95
visual-entailment-on-snli-ve-valSOHO#5Accuracy: 85.00
visual-reasoning-on-nlvr2-devSOHO#12Accuracy: 76.37
visual-reasoning-on-nlvr2-testSOHO#12Accuracy: 77.32