Teaching Large Language Models to Self-Debug

Benchmark Model Rank Results
code-generation-on-mbppGPT-4 (Self-Debugging with unit tests + trace)#15Accuracy: 80.2
code-generation-on-mbppGPT-3.5 Turbo (Self-Debugging with unit tests + trace)#17Accuracy: 72.8
code-generation-on-mbppcode-davinci-002 175B (Self-Debugging with unit tests + trace)#18Accuracy: 70.8
code-generation-on-mbppGPT-3.5 Turbo (3-shot)#24Accuracy: 67.6
code-generation-on-mbppcode-davinci-002 175B (3-shot)#35Accuracy: 61.4
code-generation-on-mbppStarCoder 15.5B (Self-Debugging with unit tests + trace)#44Accuracy: 53.2
code-generation-on-mbppStarCoder 15.5B (3-shot)#59Accuracy: 47.2