Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Benchmark Model Rank Results
video-based-generative-performanceVLM-RLAIF#2mean: 3.49Correctness of Information: 3.63Detail Orientation: 3.25