Evaluating Large Language Models Trained on Code

Benchmark Model Rank Results
code-generation-on-appsCodex 12B (Raw)#12Introductory Pass@1: 5.60%Interview Pass@1: 1.00%
multi-task-language-understanding-on-bbh-algcode-davinci-002 175B (CoT)#1Average (%): 73.9
multi-task-language-understanding-on-bbh-nlpcode-davinci-002 175B (CoT)#3Average (%): 73.5