Visual-Textual Capsule Routing for Text-Based Video Segmentation

Benchmark Model Rank Results
referring-expression-segmentation-on-a2d-sentencesVT-Capsule–AP: 0.303IoU overall: 0.568IoU mean: 0.460Precision@0.5: 0.526…
referring-expression-segmentation-on-j-hmdbVT-Capsule–AP: 0.261IoU overall: 0.535IoU mean: 0.550Precision@0.5: 0.677…