Dynamic Scene Understanding from Vision-Language Representations

Benchmark Model Rank Results
grounded-situation-recognition-on-swigOurs (CoFormer+)Top-1 Verb: 58.88Top-1 Verb & Grounded-Value: 41.28
human-object-interaction-detection-on-hicoOurs (PViC+)mAP: 46.49
situation-recognition-on-imsituOursTop-1 Verb: 58.88