ActBERT: Learning Global-Local Video-Text Representations

Benchmark Model Rank Results
action-segmentation-on-coinActBERT#7Frame accuracy: 57.0