SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

Benchmark Model Rank Results
referring-video-object-segmentation-on-long-rvosSAMWISE#3J&F: 35.6tIoU: 68.4vIoU: 28.6
referring-video-object-segmentation-on-mevisSAMWISE#7J&F: 48.3J: 45.4F: 51.2