OpenAI Introduces GeneBench-Pro for Computational Biology Evaluation
GeneBench-Pro tests AI judgment in genomics, quantitative biology, and medicine OpenAI's benchmark reveals how frontier models reason through messy science data
OpenAI has introduced GeneBenchPro, a new benchmark designed to test whether AI models can handle judgmentheavy analysis in computational biology. The benchmark focuses on tasks in genomics, quantitative biology, and translational medicine, where models must decide how to analyze messy data, revise assumptions, and determine when results are ready for use.
According to OpenAI, GeneBenchPro contains 129 synthetic problems across 10 domains and 21 subdomains. The company says the benchmark was built to reduce common evaluation flaws by using fully simulated datagenerating processes, allowing results to be graded against known targets and helping ensure that correct solutions depend on choosing the right analytical path.
OpenAI also reported results from its internal testing, saying its strongest model, GPT5.6 Sol, reached a 28.7% pass rate at the highest reasoning level, or 31.5% with Pro mode enabled. The company said the benchmark suggests frontier models are improving on complex scientific reasoning, but still fall short of reliably solving most problems.