FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
Open paper
Benchmark
Model
Rank
Results
human-judgment-correlation-on-flickr8k-expert
SoftSPICE
#2
Kendall's Tau-c: 54.2
Rank counts only results with a code link.