| QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks | ✓ Link | 98.8 | QualityFlow (Sonnet-3.5) | 2025-01-20 |
| Planning-Driven Programming: A Large Language Model Programming Workflow | ✓ Link | 98.2 | Phi-2 | 2024-11-21 |
| Introducing Claude Sonnet 4.5 | | 97.6 | Claude Sonnet 4.5 Thinking | 2025-09-29 |
| DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | ✓ Link | 97.4 | DeepSeek R1 | 2025-01-22 |
| Introducing Claude Opus 4.6 | | 97.0 | Claude Opus 4.6 Thinking | 2026-02-05 |
| GLM-5: From Vibe Coding to Agentic Engineering | ✓ Link | 97.0 | GLM 5 Thinking | 2026-02-11 |
| Execution Guided Line-by-Line Code Generation | ✓ Link | 96.95 | EG-CFG (DeepSeek-V3-0324) | 2025-06-12 |
| Introducing OpenAI o3 and o4-mini | | 96.3 | o4 Mini High | 2025-04-16 |
| Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models | ✓ Link | 95.9 | Youtu-LLM 2B | 2025-12-31 |
| Claude Opus 4.1 | | 95.7 | Claude Opus 4.1 | 2025-08-05 |
| Gemini 2.5: Our most intelligent AI model | | 95.1 | Gemini 2.5 Pro Preview 05-06 | 2025-03-25 |
| From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging | ✓ Link | 94.5 | DeepSeek-Coder-V2-Lite (MGDebugger) | 2024-10-02 |
| HyperCLOVA X THINK Technical Report | ✓ Link | 94.5 | HyperCLOVA X SEED 14B Think | 2025-06-27 |
| Introducing GPT-5.1 for developers | | 94.5 | GPT-5.1 | 2025-11-12 |
| MapCoder: Multi-Agent Code Generation for Competitive Problem Solving | ✓ Link | 93.9 | Mistral 7B | 2024-05-18 |
| DeepSeek-V3.2 Technical Report | ✓ Link | 93.9 | DeepSeek V3.2 Thinking | 2025-12-02 |
| Qwen3-Coder: Agentic Coding in the World | ✓ Link | 92.7 | Qwen3 Coder 480B A35B (exacto) | 2025-07-22 |
| Qwen3 Technical Report | ✓ Link | 92.1 | Qwen3 235B A22B Instruct 2507 | 2025-05-14 |
| ERNIE 4.5 Technical Report | ✓ Link | 92.1 | ERNIE-4.5-300B-A47B | 2025-06-30 |
| Introducing Claude 3.5 Sonnet | | 92.0 | Claude 3.5 Sonnet | 2024-06-20 |
| Introducing Claude 3.5 Sonnet | | 90.85 | Claude Sonnet 3.5 | 2024-06-20 |
| L2MAC: Large Language Model Automatic Computer for Extensive Code Generation | ✓ Link | 90.2 | L2MAC (GPT-4) | 2023-10-02 |
| Hello GPT-4o | | 90.2 | GPT-4o | 2024-05-13 |
| Mistral Large 3 | ✓ Link | 90.2 | Mistral Large 3 2512 | 2025-11-28 |
| Introducing Llama 3.1: Our most capable models to date | ✓ Link | 89.0 | Llama 3.1 (405B, Instruct) | 2024-07-23 |
| dots.llm1 Technical Report | ✓ Link | 88.4 | dots.llm1.inst | 2025-06-06 |
| Qwen2.5 Technical Report | | 87.8 | Qwen2.5-Plus | 2024-12-19 |
| Qwen2.5-VL Technical Report | ✓ Link | 87.8 | Qwen2.5-VL-72B | 2025-02-19 |
| Gemma 3 Technical Report | ✓ Link | 87.8 | Gemma 3 (27B, IT) | 2025-03-25 |
| GPT-4o mini: advancing cost-efficient intelligence | | 87.2 | GPT-4o mini | 2024-07-18 |
| DeepSeek-V3 Technical Report | ✓ Link | 87.2 | DeepSeek V3 | 2024-12-27 |
| Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance | ✓ Link | 87.2 | Falcon-H1-34B-Instruct | 2025-07-30 |
| MiniMax-01: Scaling Foundation Models with Lightning Attention | ✓ Link | 86.9 | MiniMax-Text-01 | 2025-01-14 |
| MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction | ✓ Link | 86.6 | MiniCPM-o 4.5-Instruct | 2026-04-30 |
| Qwen2 Technical Report | ✓ Link | 86.0 | Qwen2-72B-Instruct | 2024-07-15 |
| Introducing the next generation of Claude | | 84.9 | Claude 3 Opus | 2024-03-04 |
| The Llama 4 herd | ✓ Link | 84.8 | Llama 4 Maverick | 2025-04-05 |
| Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context | | 84.1 | Gemini 1.5 Pro | 2024-03-08 |
| Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step | ✓ Link | 82.9 | LLaMA 3 | 2024-02-25 |
| Command A: An Enterprise-Ready Large Language Model | ✓ Link | 82.9 | Command A | 2025-03-13 |
| InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models | ✓ Link | 82.3 | InternVL3-78B | 2025-04-14 |
| DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model | ✓ Link | 81.1 | DeepSeek-V2 Chat (RL) (0-shot) | 2024-05-07 |
| The Llama 4 herd | ✓ Link | 81.1 | Llama 4 Scout | 2025-04-05 |
| Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters | ✓ Link | 81.1 | Step-3.5-Flash Base | 2026-02-11 |
| Introducing Llama 3.1: Our most capable models to date | ✓ Link | 80.5 | Llama 3.1 (70B, Instruct) | 2024-07-23 |
| Qwen2 Technical Report | ✓ Link | 79.9 | Qwen2-57B-A14B-Instruct | 2024-07-15 |
| Qwen2.5-Omni Technical Report | ✓ Link | 78.7 | Qwen2.5-Omni-7B | 2025-03-26 |
| ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools | ✓ Link | 78.5 | GLM-4 (0520) | 2024-06-18 |