| bias-detection-on-stereoset-1 | GAL 120B | #8 | ICAT Score: 65.6LMS: 75SS: 56.2 |
| bias-detection-on-stereoset-1 | GPT-3 (text-davinci-002) | #10 | ICAT Score: 60.8LMS: 77.6SS: 60.8 |
| bias-detection-on-stereoset-1 | OPT 175B | #11 | ICAT Score: 60LMS: 74.8SS: 59.9 |
| common-sense-reasoning-on-arc-challenge | GAL 120B (zero-shot) | #9 | Accuracy: 67.9 |
| common-sense-reasoning-on-arc-challenge | GPT-3 (zero-shot) | #23 | Accuracy: 51.4 |
| common-sense-reasoning-on-arc-challenge | BLOOM (few-shot, k=5) | #37 | Accuracy: 32.9 |
| common-sense-reasoning-on-arc-challenge | OPT (few-shot, k=5) | #39 | Accuracy: 31.1 |
| common-sense-reasoning-on-arc-easy | GAL 120B (0-shot) | #8 | Accuracy: 83.8 |
| common-sense-reasoning-on-arc-easy | GPT-3 (zero-shot) | #34 | Accuracy: 68.8 |
| common-sense-reasoning-on-arc-easy | BLOOM (5-shot) | #37 | Accuracy: 40.7 |
| common-sense-reasoning-on-arc-easy | OPT (5-shot) | #39 | Accuracy: 37.4 |
| math-word-problem-solving-on-math | Minerva 540B (5-shot) mCoT | #69 | Accuracy: 33.6Parameters (Billions): 540 |
| math-word-problem-solving-on-math | GAL 120B (5-shot) mCoT | #89 | Accuracy: 20.4Parameters (Billions): 120 |
| math-word-problem-solving-on-math | GAL 120B <work> | #93 | Accuracy: 16.6Parameters (Billions): 120 |
| math-word-problem-solving-on-math | GAL 30B (5-shot) mCoT | #98 | Accuracy: 12.7Parameters (Billions): 30 |
| math-word-problem-solving-on-math | GAL 30B <work> | #100 | Accuracy: 11.4Parameters (Billions): 30 |
| math-word-problem-solving-on-math | PaLM 540B (5-shot) mCoT | #104 | Accuracy: 8.8Parameters (Billions): 540 |
| math-word-problem-solving-on-math | GPT-3 175B (8-shot) | #115 | Accuracy: 5.2Parameters (Billions): 175 |
| molecular-property-prediction-on-bace-1 | GAL 30B | #15 | ROC-AUC: 72.7 |
| molecular-property-prediction-on-bace-1 | GAL 120B | #16 | ROC-AUC: 61.7 |
| molecular-property-prediction-on-bace-1 | GAL 6.7B | #17 | ROC-AUC: 58.4 |
| molecular-property-prediction-on-bace-1 | GAL 1.3B | #18 | ROC-AUC: 57.6 |
| molecular-property-prediction-on-bace-1 | GAL 125M | #19 | ROC-AUC: 56.1 |
| molecular-property-prediction-on-bbbp-1 | Uni-Mol | #15 | ROC-AUC: 72.9 |
| molecular-property-prediction-on-bbbp-1 | GAL 120B | #23 | ROC-AUC: 66.1 |
| molecular-property-prediction-on-bbbp-1 | GAL 1.3B | #24 | ROC-AUC: 60.4 |
| molecular-property-prediction-on-bbbp-1 | GAL 30B | #25 | ROC-AUC: 59.6 |
| molecular-property-prediction-on-bbbp-1 | GAL 6.7B | #26 | ROC-AUC: 53.5 |
| molecular-property-prediction-on-bbbp-1 | GAL 125M | #27 | ROC-AUC: 39.3 |
| molecular-property-prediction-on-clintox-1 | GAL 120B | #8 | ROC-AUC: 82.6Molecules (M): 2 |
| molecular-property-prediction-on-clintox-1 | GAL 30B | #9 | ROC-AUC: 82.2Molecules (M): 2 |
| molecular-property-prediction-on-clintox-1 | GAL 6.7B | #12 | ROC-AUC: 78.4Molecules (M): 2 |
| molecular-property-prediction-on-clintox-1 | GAL 1.3B | #16 | ROC-AUC: 58.9Molecules (M): 2 |
| molecular-property-prediction-on-clintox-1 | GAL 125M | #18 | ROC-AUC: 51.8Molecules (M): 2 |
| molecular-property-prediction-on-hiv-dataset | Uni-Mol | #2 | AUC: 0.808 |
| molecular-property-prediction-on-hiv-dataset | GAL 30B | #7 | AUC: 0.759 |
| molecular-property-prediction-on-hiv-dataset | GAL 120B | #8 | AUC: 0.745 |
| molecular-property-prediction-on-hiv-dataset | GAL 1.3B | #9 | AUC: 0.724 |
| molecular-property-prediction-on-hiv-dataset | GAL 6.7B | #10 | AUC: 0.722 |
| molecular-property-prediction-on-hiv-dataset | GAL 125M | #11 | AUC: 0.702 |
| molecular-property-prediction-on-moleculenet | Uni-Mol | #1 | AUC: 0.77 |
| molecular-property-prediction-on-moleculenet | GAL 30B | #2 | AUC: 0.69 |
| molecular-property-prediction-on-moleculenet | GAL 6.7B | #3 | AUC: 0.64 |
| molecular-property-prediction-on-moleculenet | GAL 1.3B | #4 | AUC: 0.619 |
| molecular-property-prediction-on-moleculenet | GAL 125M | #5 | AUC: 0.581 |
| molecular-property-prediction-on-sider-1 | GAL 120B | #11 | ROC-AUC: 63.2 |
| molecular-property-prediction-on-sider-1 | GAL 30B | #13 | ROC-AUC: 61.3 |
| molecular-property-prediction-on-sider-1 | GAL 125M | #15 | ROC-AUC: 55.9 |
| molecular-property-prediction-on-sider-1 | GAL 6.7B | #16 | ROC-AUC: 55.9 |
| molecular-property-prediction-on-sider-1 | GAL 1.3B | #17 | ROC-AUC: 54.0 |
| molecular-property-prediction-on-tox21-1 | Uni-Mol | #5 | ROC-AUC: 79.6 |
| molecular-property-prediction-on-tox21-1 | GAL 120B | #14 | ROC-AUC: 68.9 |
| molecular-property-prediction-on-tox21-1 | GAL 30B | #15 | ROC-AUC: 68.5 |
| molecular-property-prediction-on-tox21-1 | GAL 6.7B | #16 | ROC-AUC: 63.9 |
| molecular-property-prediction-on-tox21-1 | GAL 1.3B | #17 | ROC-AUC: 60.6 |
| molecular-property-prediction-on-tox21-1 | GAL 125M | #18 | ROC-AUC: 54.3 |
| multi-task-language-understanding-on-mmlu | GAL 120B (zero-shot) | #18 | Average (%): 52.6 |
| multiple-choice-question-answering-mcqa-on-10 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 41.5 |
| multiple-choice-question-answering-mcqa-on-10 | GAL 120B (zero-shot) | #2 | Accuracy: 38.1 |
| multiple-choice-question-answering-mcqa-on-10 | Gopher (few-shot, k=5) | #3 | Accuracy: 33.6 |
| multiple-choice-question-answering-mcqa-on-10 | BLOOM (few-shot, k=5) | #4 | Accuracy: 27.6 |
| multiple-choice-question-answering-mcqa-on-10 | OPT (few-shot, k=5) | #5 | Accuracy: 25.7 |
| multiple-choice-question-answering-mcqa-on-11 | Chinchilla (few-shot, k=5) | #4 | Accuracy: 79.9 |
| multiple-choice-question-answering-mcqa-on-11 | Gopher (few-shot, k=5) | #5 | Accuracy: 70.8 |
| multiple-choice-question-answering-mcqa-on-11 | GAL 120B (zero-shot) | #6 | Accuracy: 68.8 |
| multiple-choice-question-answering-mcqa-on-11 | OPT (few-shot, k=5) | #7 | Accuracy: 30.6 |
| multiple-choice-question-answering-mcqa-on-11 | BLOOM (few-shot, k=5) | #8 | Accuracy: 28.5 |
| multiple-choice-question-answering-mcqa-on-12 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 80.3 |
| multiple-choice-question-answering-mcqa-on-12 | Gopher (few-shot, k=5) | #2 | Accuracy: 71.3 |
| multiple-choice-question-answering-mcqa-on-12 | GAL 120B (zero-shot) | #3 | Accuracy: 69.4 |
| multiple-choice-question-answering-mcqa-on-12 | BLOOM (few-shot, k=5) | #4 | Accuracy: 29.4 |
| multiple-choice-question-answering-mcqa-on-12 | OPT (few-shot, k=5) | #5 | Accuracy: 27.7 |
| multiple-choice-question-answering-mcqa-on-13 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 51 |
| multiple-choice-question-answering-mcqa-on-13 | GAL 120B (zero-shot) | #2 | Accuracy: 46 |
| multiple-choice-question-answering-mcqa-on-13 | Gopher (few-shot, k=5) | #3 | Accuracy: 45 |
| multiple-choice-question-answering-mcqa-on-13 | OPT (few-shot, k=5) | #4 | Accuracy: 30 |
| multiple-choice-question-answering-mcqa-on-13 | BLOOM (few-shot, k=5) | #5 | Accuracy: 19 |
| multiple-choice-question-answering-mcqa-on-14 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 58.1 |
| multiple-choice-question-answering-mcqa-on-14 | GAL 120B (zero-shot) | #2 | Accuracy: 47.8 |
| multiple-choice-question-answering-mcqa-on-14 | BLOOM (few-shot, k=5) | #3 | Accuracy: 23.2 |
| multiple-choice-question-answering-mcqa-on-14 | OPT (few-shot, k=5) | #4 | Accuracy: 21.7 |
| multiple-choice-question-answering-mcqa-on-15 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 51.0 |
| multiple-choice-question-answering-mcqa-on-15 | GAL 120B (zero-shot) | #2 | Accuracy: 49 |
| multiple-choice-question-answering-mcqa-on-15 | OPT (few-shot, k=5) | #3 | Accuracy: 17.0 |
| multiple-choice-question-answering-mcqa-on-15 | BLOOM (few-shot, k=5) | #4 | Accuracy: 6.0 |
| multiple-choice-question-answering-mcqa-on-16 | GAL 120B (zero-shot) | #1 | Accuracy: 32.6 |
| multiple-choice-question-answering-mcqa-on-16 | Chinchilla (few-shot, k=5) | #2 | Accuracy: 31.9 |
| multiple-choice-question-answering-mcqa-on-16 | BLOOM (few-shot, k=5) | #3 | Accuracy: 27 |
| multiple-choice-question-answering-mcqa-on-16 | OPT (few-shot, k=5) | #4 | Accuracy: 24.4 |
| multiple-choice-question-answering-mcqa-on-16 | Gopher (few-shot, k=5) | #5 | Accuracy: 23.7 |
| multiple-choice-question-answering-mcqa-on-17 | GAL 120B (zero-shot) | #1 | Accuracy: 62.8 |
| multiple-choice-question-answering-mcqa-on-17 | Chinchilla (few-shot, k=5) | #2 | Accuracy: 62.1 |
| multiple-choice-question-answering-mcqa-on-17 | Gopher (few-shot, k=5) | #3 | Accuracy: 60 |
| multiple-choice-question-answering-mcqa-on-17 | OPT (few-shot, k=5) | #4 | Accuracy: 36.6 |
| multiple-choice-question-answering-mcqa-on-17 | BLOOM (few-shot, k=5) | #5 | Accuracy: 32.4 |
| multiple-choice-question-answering-mcqa-on-18 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 46.1 |
| multiple-choice-question-answering-mcqa-on-18 | GAL 120B (zero-shot) | #2 | Accuracy: 42.2 |
| multiple-choice-question-answering-mcqa-on-18 | Gopher (few-shot, k=5) | #3 | Accuracy: 34.3 |
| multiple-choice-question-answering-mcqa-on-18 | OPT (few-shot, k=5) | #4 | Accuracy: 21.6 |
| multiple-choice-question-answering-mcqa-on-18 | BLOOM (few-shot, k=5) | #5 | Accuracy: 18.6 |
| multiple-choice-question-answering-mcqa-on-19 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 36.4 |
| multiple-choice-question-answering-mcqa-on-19 | GAL 120B (zero-shot) | #2 | Accuracy: 33.8 |
| multiple-choice-question-answering-mcqa-on-19 | OPT (few-shot, k=5) | #3 | Accuracy: 29.8 |
| multiple-choice-question-answering-mcqa-on-19 | BLOOM (few-shot, k=5) | #4 | Accuracy: 25.2 |
| multiple-choice-question-answering-mcqa-on-2 | Gopher (few-shot, k=5) | #1 | Accuracy: 35.7 |
| multiple-choice-question-answering-mcqa-on-2 | Chinchilla (few-shot, k=5) | #2 | Accuracy: 33.3 |
| multiple-choice-question-answering-mcqa-on-2 | GAL 120B (zero-shot) | #3 | Accuracy: 32.5 |
| multiple-choice-question-answering-mcqa-on-2 | OPT (few-shot, k=5) | #4 | Accuracy: 29.4 |
| multiple-choice-question-answering-mcqa-on-2 | BLOOM (few-shot, k=5) | #5 | Accuracy: 26.2 |
| multiple-choice-question-answering-mcqa-on-20 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 58.8 |
| multiple-choice-question-answering-mcqa-on-20 | Gopher (few-shot, k=5) | #2 | Accuracy: 50 |
| multiple-choice-question-answering-mcqa-on-20 | OPT (few-shot, k=5) | #3 | Accuracy: 43.5 |
| multiple-choice-question-answering-mcqa-on-20 | GAL 120B (zero-shot) | #4 | Accuracy: 41.2 |
| multiple-choice-question-answering-mcqa-on-20 | BLOOM (few-shot, k=5) | #5 | Accuracy: 19.4 |
| multiple-choice-question-answering-mcqa-on-21 | GAL 120B (zero-shot) | #16 | Dev Set (Acc-%): 0.529 |
| multiple-choice-question-answering-mcqa-on-21 | BLOOM (few-shot, k=5) | #20 | Dev Set (Acc-%): 0.325 |
| multiple-choice-question-answering-mcqa-on-21 | OPT (few-shot, k=5) | #21 | Dev Set (Acc-%): 0.296 |
| multiple-choice-question-answering-mcqa-on-3 | GAL 30B (zero-shot) | #1 | Accuracy: 33.3 |
| multiple-choice-question-answering-mcqa-on-3 | Chinchilla (few-shot, k=5) | #2 | Accuracy: 31 |
| multiple-choice-question-answering-mcqa-on-3 | GAL 120B (zero-shot) | #3 | Accuracy: 27 |
| multiple-choice-question-answering-mcqa-on-3 | Gopher (few-shot, k=5) | #4 | Accuracy: 25 |
| multiple-choice-question-answering-mcqa-on-3 | OPT (few-shot, k=5) | #5 | Accuracy: 21 |
| multiple-choice-question-answering-mcqa-on-4 | Gopher (few-shot, k=5) | #1 | Accuracy: 43 |
| multiple-choice-question-answering-mcqa-on-4 | GAL 120B (zero-shot) | #2 | Accuracy: 42.1 |
| multiple-choice-question-answering-mcqa-on-4 | Chinchilla (few-shot, k=5) | #3 | Accuracy: 38.6 |
| multiple-choice-question-answering-mcqa-on-4 | BLOOM (few-shot, k=5) | #4 | Accuracy: 23.7 |
| multiple-choice-question-answering-mcqa-on-4 | OPT (few-shot, k=5) | #5 | Accuracy: 21 |
| multiple-choice-question-answering-mcqa-on-5 | GAL 120B (zero-shot) | #1 | Accuracy: 70 |
| multiple-choice-question-answering-mcqa-on-5 | Chinchilla (few-shot, k=5) | #2 | Accuracy: 58 |
| multiple-choice-question-answering-mcqa-on-5 | Gopher (few-shot, k=5) | #3 | Accuracy: 54 |
| multiple-choice-question-answering-mcqa-on-5 | OPT (few-shot, k=5) | #4 | Accuracy: 30 |
| multiple-choice-question-answering-mcqa-on-5 | BLOOM (few-shot, k=5) | #5 | Accuracy: 25 |
| multiple-choice-question-answering-mcqa-on-6 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 41.1 |
| multiple-choice-question-answering-mcqa-on-6 | GAL 120B (zero-shot) | #2 | Accuracy: 38.4 |
| multiple-choice-question-answering-mcqa-on-6 | OPT (few-shot, k=5) | #3 | Accuracy: 28.6 |
| multiple-choice-question-answering-mcqa-on-6 | BLOOM (few-shot, k=5) | #4 | Accuracy: 25 |
| multiple-choice-question-answering-mcqa-on-7 | GAL 120B (zero-shot) | #1 | Accuracy: 43 |
| multiple-choice-question-answering-mcqa-on-7 | Gopher (few-shot, k=5) | #2 | Accuracy: 37 |
| multiple-choice-question-answering-mcqa-on-7 | OPT (few-shot, k=5) | #3 | Accuracy: 33 |
| multiple-choice-question-answering-mcqa-on-7 | Chinchilla (few-shot, k=5) | #4 | Accuracy: 32 |
| multiple-choice-question-answering-mcqa-on-7 | BLOOM (few-shot, k=5) | #5 | Accuracy: 25 |
| multiple-choice-question-answering-mcqa-on-8 | GAL 30B (zero-shot) | #4 | Accuracy: 70 |
| multiple-choice-question-answering-mcqa-on-8 | Chinchilla (few-shot, k=5) | #5 | Accuracy: 69 |
| multiple-choice-question-answering-mcqa-on-8 | GAL 120B (zero-shot) | #6 | Accuracy: 68 |
| multiple-choice-question-answering-mcqa-on-8 | BLOOM (few-shot, k=5) | #7 | Accuracy: 36 |
| multiple-choice-question-answering-mcqa-on-8 | OPT (few-shot, k=5) | #8 | Accuracy: 35 |
| multiple-choice-question-answering-mcqa-on-9 | Chinchilla (few-shot, k=5) | #1 | Accuracy: 73.0 |
| multiple-choice-question-answering-mcqa-on-9 | Gopher (few-shot, k=5) | #2 | Accuracy: 65.8 |
| multiple-choice-question-answering-mcqa-on-9 | GAL 120B (zero-shot) | #3 | Accuracy: 65.1 |
| multiple-choice-question-answering-mcqa-on-9 | BLOOM (few-shot, k=5) | #4 | Accuracy: 25.7 |
| multiple-choice-question-answering-mcqa-on-9 | OPT (few-shot, k=5) | #5 | Accuracy: 23.0 |
| protein-function-prediction-on-caspsimseq | GAL 120B | #1 | ROUGE-L: 0.252 |
| protein-function-prediction-on-caspsimseq | GAL 30B | #2 | ROUGE-L: 0.137 |
| protein-function-prediction-on-caspsimseq | GAL 6.7B | #3 | ROUGE-L: 0.109 |
| protein-function-prediction-on-caspsimseq | GAL 1.3B | #4 | ROUGE-L: 0.069 |
| protein-function-prediction-on-caspsimseq | GAL 125M | #5 | ROUGE-L: 0.062 |
| protein-function-prediction-on-paenseq | GAL 120B | #1 | ROUGE-L: 0.272 |
| protein-function-prediction-on-paenseq | GAL 30B | #2 | ROUGE-L: 0.196 |
| protein-function-prediction-on-paenseq | GAL 6.7B | #3 | ROUGE-L: 0.137 |
| protein-function-prediction-on-paenseq | GAL 1.3B | #4 | ROUGE-L: 0.084 |
| protein-function-prediction-on-paenseq | GAL 125M | #5 | ROUGE-L: 0.073 |
| protein-function-prediction-on-uniprotseq | GAL 120B | #1 | ROUGE-L: 0.252 |
| protein-function-prediction-on-uniprotseq | GAL 30B | #2 | ROUGE-L: 0.186 |
| protein-function-prediction-on-uniprotseq | GAL 6.7B | #3 | ROUGE-L: 0.111 |
| protein-function-prediction-on-uniprotseq | GAL 1.3B | #4 | ROUGE-L: 0.079 |
| protein-function-prediction-on-uniprotseq | GAL 125M | #5 | ROUGE-L: 0.061 |
| protein-structure-prediction-on-caspseq | GAL 120B | #1 | Validation perplexity: 17.26 |
| protein-structure-prediction-on-caspseq | GAL 30B | #2 | Validation perplexity: 17.27 |
| protein-structure-prediction-on-caspseq | GAL 6.7B | #3 | Validation perplexity: 17.29 |
| protein-structure-prediction-on-caspseq | GAL 1.3B | #4 | Validation perplexity: 17.58 |
| protein-structure-prediction-on-caspseq | GAL 125M | #5 | Validation perplexity: 20.62 |
| protein-structure-prediction-on-caspsimseq | GAL 120B | #1 | Validation perplexity: 12.77 |
| protein-structure-prediction-on-caspsimseq | GAL 30B | #2 | Validation perplexity: 15.42 |
| protein-structure-prediction-on-caspsimseq | GAL 6.7B | #3 | Validation perplexity: 16.35 |
| protein-structure-prediction-on-caspsimseq | GAL 1.3B | #4 | Validation perplexity: 17.04 |
| protein-structure-prediction-on-caspsimseq | GAL 125M | #5 | Validation perplexity: 19.18 |
| protein-structure-prediction-on-paenseq | GAL 120B | #1 | Validation perplexity: 3.14 |
| protein-structure-prediction-on-paenseq | GAL 30B | #2 | Validation perplexity: 4.28 |
| protein-structure-prediction-on-paenseq | GAL 6.7B | #3 | Validation perplexity: 7.76 |
| protein-structure-prediction-on-paenseq | GAL 1.3B | #4 | Validation perplexity: 12.53 |
| protein-structure-prediction-on-paenseq | GAL 125M | #5 | Validation perplexity: 16.35 |
| protein-structure-prediction-on-uniprotseq | GAL 120B | #1 | Validation perplexity: 5.54 |
| protein-structure-prediction-on-uniprotseq | GAL 30B | #2 | Validation perplexity: 8.23 |
| protein-structure-prediction-on-uniprotseq | GAL 6.7B | #3 | Validation perplexity: 11.58 |
| protein-structure-prediction-on-uniprotseq | GAL 1.3B | #4 | Validation perplexity: 15.82 |
| protein-structure-prediction-on-uniprotseq | GAL 125M | #5 | Validation perplexity: 19.05 |
| question-answering-on-bioasq | GAL 120B (zero-shot) | #2 | Accuracy: 94.3 |
| question-answering-on-bioasq | BLOOM (zero-shot) | #4 | Accuracy: 91.4 |
| question-answering-on-bioasq | OPT (zero-shot) | #6 | Accuracy: 81.4 |
| question-answering-on-medqa-usmle | GAL 120B (zero-shot) | #16 | Accuracy: 44.4 |
| question-answering-on-medqa-usmle | BLOOM (few-shot, k=5) | #21 | Accuracy: 23.3 |
| question-answering-on-medqa-usmle | OPT (few-shot, k=5) | #22 | Accuracy: 22.8 |
| question-answering-on-pubmedqa | GAL 120B (zero-shot) | #9 | Accuracy: 77.6 |
| question-answering-on-pubmedqa | BLOOM (zero-shot) | #15 | Accuracy: 73.6 |
| question-answering-on-pubmedqa | OPT (zero-shot) | #19 | Accuracy: 70.2 |
| question-answering-on-truthfulqa | GAL 120B | #7 | MC1: 0.26 |
| question-answering-on-truthfulqa | GAL 30B | #9 | MC1: 0.24 |
| question-answering-on-truthfulqa | OPT 175B | #15 | MC1: 0.21 |
| question-answering-on-truthfulqa | GAL 125M | #18 | MC1: 0.19 |
| question-answering-on-truthfulqa | GAL 1.3B | #19 | MC1: 0.19 |
| question-answering-on-truthfulqa | GAL 6.7B | #20 | MC1: 0.19 |
| stereotypical-bias-analysis-on-crows-pairs | GAL 120B | #1 | Gender: 51.9Religion: 51.9Race/Color: 59.9Sexual Orientation: 77.4… |
| tdc-admet-benchmarking-group-on-tdcommons | Galactica-GAL-120B | #8 | TDC.BBB_Martins: 0.661 |
| tdc-admet-benchmarking-group-on-tdcommons | Galactica-GAL-1.3B | #9 | TDC.BBB_Martins: 0.604 |
| tdc-admet-benchmarking-group-on-tdcommons | Galactica-GAL-30B | #10 | TDC.BBB_Martins: 0.596 |
| tdc-admet-benchmarking-group-on-tdcommons | Galactica-GAL-6.7B | #11 | TDC.BBB_Martins: 0.535 |
| tdc-admet-benchmarking-group-on-tdcommons | Galactica-GAL-125M | #12 | TDC.BBB_Martins: 0.393 |
| word-sense-disambiguation-on-big-bench | OPT 175B | #3 | Accuracy: 49.1 |
| word-sense-disambiguation-on-big-bench | GAL 120B (few-shot, k=5) | #4 | Accuracy: 48.7 |
| word-sense-disambiguation-on-big-bench | GAL 30B (few-shot, k=5) | #5 | Accuracy: 47.0 |
| word-sense-disambiguation-on-big-bench | BLOOM 176B | #6 | Accuracy: 1.3 |