Training Complex Models with Multi-Task Weak Supervision

Benchmark Model Rank Results
natural-language-inference-on-multinliSnorkel MeTaL (ensemble)#16Matched: 87.6Mismatched: 87.2
paraphrase-identification-on-quora-question-pairsSnorkel MeTaL(ensemble)#8F1: 73.1Accuracy: 89.9
sentiment-analysis-on-sst-2-binary-classificationSnorkel MeTaL(ensemble)#18Accuracy: 96.2