LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Benchmark Model Rank Results
document-image-classification-on-rvl-cdipLayoutLMv2LARGE#4Accuracy: 95.64%
document-image-classification-on-rvl-cdipLayoutLMv2BASE#9Accuracy: 95.25%Parameters: 200M
key-information-extraction-on-cordLayoutLMv2LARGE#6F1: 96.01
key-information-extraction-on-cordLayoutLMv2BASE#7F1: 94.95
key-information-extraction-on-sroieLayoutLMv2LARGE (Excluding OCR mismatch)#1F1: 97.81
key-information-extraction-on-sroieLayoutLMv2LARGE#3F1: 96.61
key-information-extraction-on-sroieLayoutLMv2BASE#4F1: 96.25
key-value-pair-extraction-on-rfund-enLayoutLMv2_base#12key-value pair F1: 49.06
relation-extraction-on-funsdLayoutLMv2 large#7F1: 70.57
semantic-entity-labeling-on-funsdLayoutLMv2LARGE#10F1: 84.2
semantic-entity-labeling-on-funsdLayoutLMv2BASE#11F1: 82.76
visual-question-answering-on-docvqa-testLayoutLMv2LARGE#13ANLS: 0.8672
visual-question-answering-on-docvqa-testLayoutLMv2BASE#21ANLS: 0.7808