Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

Benchmark Model Rank Results
referring-video-object-segmentation-on-ref-davis17HCD–J&F: 71.2J: 67.5F: 75.0