Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Open paper
Benchmark
Model
Rank
Results
video-based-generative-performance
VLM-RLAIF
#2
mean: 3.49
Correctness of Information: 3.63
Detail Orientation: 3.25
…
Rank counts only results with a code link.