| Benchmark | Model | Rank | Results |
|---|---|---|---|
| action-segmentation-on-coin | VLM | #5 | Frame accuracy: 68.4 |
| temporal-action-localization-on-crosstask | VLM | #2 | Recall: 46.5 |
| video-captioning-on-youcook2 | VLM | #4 | BLEU-4: 12.27BLEU-3: 17.78CIDEr: 1.3869ROUGE-L: 41.51… |
| video-retrieval-on-msr-vtt-1ka | VLM | #43 | text-to-video R@1: 28.10text-to-video R@5: 55.50… |
| video-retrieval-on-youcook2 | VLM | #5 | text-to-video R@1: 27.05text-to-video R@5: 56.88… |