OpenCodePapers

code-generation-on-humaneval

Code Generation
Dataset Link
Results over time
Click legend items to toggle metrics. Hover points for model names.
Leaderboard
PaperCodePass@1ModelNameReleaseDate
QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks✓ Link98.8QualityFlow (Sonnet-3.5)2025-01-20
Planning-Driven Programming: A Large Language Model Programming Workflow✓ Link98.2Phi-22024-11-21
Introducing Claude Sonnet 4.597.6Claude Sonnet 4.5 Thinking2025-09-29
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning✓ Link97.4DeepSeek R12025-01-22
Introducing Claude Opus 4.697.0Claude Opus 4.6 Thinking2026-02-05
GLM-5: From Vibe Coding to Agentic Engineering✓ Link97.0GLM 5 Thinking2026-02-11
Execution Guided Line-by-Line Code Generation✓ Link96.95EG-CFG (DeepSeek-V3-0324)2025-06-12
Introducing OpenAI o3 and o4-mini96.3o4 Mini High2025-04-16
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models✓ Link95.9Youtu-LLM 2B2025-12-31
Claude Opus 4.195.7Claude Opus 4.12025-08-05
Gemini 2.5: Our most intelligent AI model95.1Gemini 2.5 Pro Preview 05-062025-03-25
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging✓ Link94.5DeepSeek-Coder-V2-Lite (MGDebugger)2024-10-02
HyperCLOVA X THINK Technical Report✓ Link94.5HyperCLOVA X SEED 14B Think2025-06-27
Introducing GPT-5.1 for developers94.5GPT-5.12025-11-12
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving✓ Link93.9Mistral 7B2024-05-18
DeepSeek-V3.2 Technical Report✓ Link93.9DeepSeek V3.2 Thinking2025-12-02
Qwen3-Coder: Agentic Coding in the World✓ Link92.7Qwen3 Coder 480B A35B (exacto)2025-07-22
Qwen3 Technical Report✓ Link92.1Qwen3 235B A22B Instruct 25072025-05-14
ERNIE 4.5 Technical Report✓ Link92.1ERNIE-4.5-300B-A47B2025-06-30
Introducing Claude 3.5 Sonnet92.0Claude 3.5 Sonnet2024-06-20
Introducing Claude 3.5 Sonnet90.85Claude Sonnet 3.52024-06-20
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation✓ Link90.2L2MAC (GPT-4)2023-10-02
Hello GPT-4o90.2GPT-4o2024-05-13
Mistral Large 3✓ Link90.2Mistral Large 3 25122025-11-28
Introducing Llama 3.1: Our most capable models to date✓ Link89.0Llama 3.1 (405B, Instruct)2024-07-23
dots.llm1 Technical Report✓ Link88.4dots.llm1.inst2025-06-06
Qwen2.5 Technical Report87.8Qwen2.5-Plus2024-12-19
Qwen2.5-VL Technical Report✓ Link87.8Qwen2.5-VL-72B2025-02-19
Gemma 3 Technical Report✓ Link87.8Gemma 3 (27B, IT)2025-03-25
GPT-4o mini: advancing cost-efficient intelligence87.2GPT-4o mini2024-07-18
DeepSeek-V3 Technical Report✓ Link87.2DeepSeek V32024-12-27
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance✓ Link87.2Falcon-H1-34B-Instruct2025-07-30
MiniMax-01: Scaling Foundation Models with Lightning Attention✓ Link86.9MiniMax-Text-012025-01-14
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction✓ Link86.6MiniCPM-o 4.5-Instruct2026-04-30
Qwen2 Technical Report✓ Link86.0Qwen2-72B-Instruct2024-07-15
Introducing the next generation of Claude84.9Claude 3 Opus2024-03-04
The Llama 4 herd✓ Link84.8Llama 4 Maverick2025-04-05
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context84.1Gemini 1.5 Pro2024-03-08
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step✓ Link82.9LLaMA 32024-02-25
Command A: An Enterprise-Ready Large Language Model✓ Link82.9Command A2025-03-13
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models✓ Link82.3InternVL3-78B2025-04-14
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model✓ Link81.1DeepSeek-V2 Chat (RL) (0-shot)2024-05-07
The Llama 4 herd✓ Link81.1Llama 4 Scout2025-04-05
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters✓ Link81.1Step-3.5-Flash Base2026-02-11
Introducing Llama 3.1: Our most capable models to date✓ Link80.5Llama 3.1 (70B, Instruct)2024-07-23
Qwen2 Technical Report✓ Link79.9Qwen2-57B-A14B-Instruct2024-07-15
Qwen2.5-Omni Technical Report✓ Link78.7Qwen2.5-Omni-7B2025-03-26
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools✓ Link78.5GLM-4 (0520)2024-06-18