You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Benchmark Model Rank Results
action-detection-on-j-hmdbYOWO + LFB#3Frame-mAP 0.5: 75.7Video-mAP 0.2: 88.3Video-mAP 0.5: 85.9
action-detection-on-j-hmdbYOWO#4Frame-mAP 0.5: 74.4Video-mAP 0.2: 87.8Video-mAP 0.5: 85.7
action-detection-on-ucf101-24YOWO + LFB#2Frame-mAP 0.5: 87.3Video-mAP 0.1: 86.1Video-mAP 0.2: 78.6
action-detection-on-ucf101-24YOWO#4Frame-mAP 0.5: 80.4Video-mAP 0.1: 82.5Video-mAP 0.2: 75.8