Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition

Benchmark Model Rank Results
action-classification-on-charadesAdaFocus (weak supervision, MViT-B-24, 32x3)MAP: 47.8
action-classification-on-charadesAdaFocus (weak supervision, MViT-B-K400-pretrain, 16x4)MAP: 41.4
action-classification-on-charadesAdaFocus (weak supervision, Slowfast-R50, 16x8)MAP: 39.3
action-classification-on-charadesAdaFocus (weak supervision, X3D-L, 32x3)MAP: 41.2
action-segmentation-on-breakfast-1AdaFocus (newly extracted I3D-features, LT-Context model)Average F1: 76.2F1@50%: 67.5F1@25%: 79.0F1@10%: 82.1Edit: 78.3
long-video-activity-recognition-on-breakfastAdaFocus (I3D-Breakfast-Pretrain-feature, GHRM)mAP: 69.6
long-video-activity-recognition-on-breakfastAdaFocus (I3D-Breakfast-Pretrain-feature, Timeception)mAP: 70.4
long-video-activity-recognition-on-breakfastAdaFocus (MViT-Breakfast-Pretrain-feature, GHRM)mAP: 79.5
long-video-activity-recognition-on-breakfastAdaFocus (MViT-Breakfast-Pretrain-feature, Timeception)mAP: 79.2
temporal-sentence-grounding-on-charades-staAdaFocus (Full, I3D-Charades-Pretrain-feature, MMN model)R1@0.7: 35.6R1@0.5: 56.7R5@0.7: 65.0R5@0.5: 87.9
temporal-sentence-grounding-on-charades-staAdaFocus (Full, MViT-Charades-Pretrain-feature, MMN model)R1@0.7: 38.6R1@0.5: 62.4R5@0.7: 66.4R5@0.5: 89.4
temporal-sentence-grounding-on-charades-staAdaFocus (Semi-weak, I3D-Charades-Pretrain-feature, D3G model)R1@0.7: 21.1R1@0.5: 46.9R5@0.7: 49.2R5@0.5: 79.3
temporal-sentence-grounding-on-charades-staAdaFocus (Semi-weak, MViT-Charades-Pretrain-feature, D3G model)R1@0.7: 21.8R1@0.5: 50.1R5@0.7: 54.6R5@0.5: 86.1
temporal-sentence-grounding-on-charades-staAdaFocus (Weak, I3D-Charades-Pretrain-feature, CPL model)R1@0.7: 22.4R1@0.5: 49.1R5@0.7: 51.8R5@0.5: 84.2
temporal-sentence-grounding-on-charades-staAdaFocus (Weak, MViT-Charades-Pretrain-feature, CPL model)R1@0.7: 23.2R1@0.5: 51.7R5@0.7: 52.6R5@0.5: 85.2