FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Benchmark Model Rank Results
human-judgment-correlation-on-flickr8k-expertSoftSPICE#2Kendall's Tau-c: 54.2