Llama 2: Open Foundation and Fine-Tuned Chat Models

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kLLaMA 2 70B (on-shot)#87Accuracy: 56.8Parameters (Billion): 70
code-generation-on-mbppLlama 2 70B (zero-shot)#65Accuracy: 45
code-generation-on-mbppLlama 2 34B (0-shot)#78Accuracy: 33
code-generation-on-mbppLlama 2 13B (0-shot)#79Accuracy: 30.6
code-generation-on-mbppLlama 2 7B (0-shot)#86Accuracy: 20.8
math-word-problem-solving-on-mawpsLLaMA 2-Chat#15Accuracy (%): 82.4
math-word-problem-solving-on-svampLLaMA 2-Chat#11Execution Accuracy: 69.2
multi-task-language-understanding-on-mmluLLaMA 2 34B (5-shot)#11Average (%): 62.6
multi-task-language-understanding-on-mmluLLaMA 2 13B (5-shot)#16Average (%): 54.8
multi-task-language-understanding-on-mmluLLaMA 2 7B (5-shot)#21Average (%): 45.3
multiple-choice-question-answering-mcqa-on-25Llama2-7B#5Accuracy: 43.38
multiple-choice-question-answering-mcqa-on-25Llama2-7B-chat#6Accuracy: 40.07
question-answering-on-boolqLLaMA 2 70B (0-shot)#16Accuracy: 85
question-answering-on-boolqLLaMA 2 34B (0-shot)#20Accuracy: 83.7
question-answering-on-boolqLLaMA 2 13B (0-shot)#23Accuracy: 81.7
question-answering-on-boolqLLaMA 2 7B (zero-shot)#28Accuracy: 77.4
question-answering-on-multitqLLaMA2#5Hits@1: 18.5
question-answering-on-natural-questionsLLaMA 2 70B (one-shot)#24EM: 33.0
question-answering-on-piqaLLaMA 2 70B (0-shot)#17Accuracy: 82.8
question-answering-on-piqaLLaMA 2 34B (0-shot)#23Accuracy: 81.9
question-answering-on-piqaLLaMA 2 13B (0-shot)#31Accuracy: 80.5
question-answering-on-piqaLLaMA 2 7B (0-shot)#37Accuracy: 78.8
question-answering-on-triviaqaLLaMA 2 70B (one-shot)#2EM: 85
sentence-completion-on-hellaswagLLaMA 2 70B (0-shot)#20Accuracy: 85.3
sentence-completion-on-hellaswagLLaMA 2 34B (0-shot)#26Accuracy: 83.3
sentence-completion-on-hellaswagLLaMA 2 13B (0-shot)#34Accuracy: 80.7
sentence-completion-on-hellaswagLLaMA 2 7B (0-shot)#40Accuracy: 77.2