SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents
Open paper
Benchmark
Model
Rank
Results
code-generation-on-mbpp
GPT-3.5 Turbo + FlowGenScrum + Test
–
Accuracy: 83.8±0.6
Rank counts only results with a code link.