OpenAI has launched GeneBench-Pro, a new benchmark designed to assess AI’s capabilities in making complex analytical decisions in computational biology. The benchmark consists of 129 artificially created problems across areas such as genomics and quantitative biology. Notably, OpenAI’s latest model, GPT-5.6 Sol, achieved a success rate of 28.7% on these problems at its highest reasoning level, a significant improvement from under 5% with its predecessor, GPT-5. As part of its commitment to open evaluation, OpenAI is releasing 10 of the benchmark questions on Hugging Face and providing a 50-question subset to Artificial Analysis for independent testing.

OpenAI: OpenAI develops and deploys advanced artificial intelligence systems focused on research and practical applications. The company released GeneBench-Pro to test AI performance on complex analytical tasks in specialized scientific domains.
GPT-5.6 Sol: GPT-5.6 Sol is OpenAI’s leading model designed for high-level reasoning tasks. It was evaluated on the newly released GeneBench-Pro benchmark to demonstrate progress in computational biology applications.
GeneBench-Pro: GeneBench-Pro is a research benchmark created to evaluate whether AI agents can handle judgment-intensive decisions across genomics, quantitative biology, and translational medicine. OpenAI introduced it publicly and is making portions available for community and third-party use.

Benchmark Release: OpenAI is open-sourcing a selection of the benchmark questions to support broader evaluation efforts.
Third-Party Evaluation: A subset of the benchmark has been shared with independent platforms to enable standardized external testing.