OpenAI audits coding benchmark and finds flaws

SWE-bench Pro audit shows 30% of tasks are broken, skewing coding results. See why OpenAI withdrew its recommendation and what it means for model evals.

OpenAI said it audited SWEBench Pro to check whether the benchmark still provides a reliable measure of coding ability. The company said accurate evaluations matter for deployment and safety decisions, and that flawed tests can distort how model capabilities are understood. In its review, OpenAI said it found evidence that a substantial share of tasks in the dataset were broken. The company said the main issues included overly strict tests, underspecified prompts, lowcoverage tests, and prompts that could mislead models. OpenAI said it used an automated pipeline, investigator agents, and human reviewers to assess flagged tasks. Based on those checks, it estimated that about 30% of SWEbench Pro tasks were broken and advised model developers to examine the benchmark results carefully. It also said it is withdrawing its earlier recommendation to switch to SWEBench Pro.