Measuring Coding Challenge Competence With APPS
Open paper
Benchmark
Model
Rank
Results
code-generation-on-apps
GPT-Neo 2.7B
#14
Introductory Pass@1: 3.90%
Interview Pass@1: 0.57%
…
Rank counts only results with a code link.