In Q2 2026, analysis of contributions to OpenAI’s Codex repository revealed that 8% of contributor-days involved estimated work exceeding 24 hours, indicating a notable increase from 2% in Q2 2025. This uptick suggests a growing efficiency among AI engineers, potentially influenced by the integration of AI coding assistants in leading AI labs, which has been associated with faster iteration cycles in both routine and complex tasks. The assessment utilized a methodology that involved AI models estimating the time required for experienced engineers to replicate code changes, highlighting the evolving productivity dynamics in software development.

METR: METR is a research organization that develops methodologies for evaluating AI capabilities through time-estimation benchmarks. Its approach to assessing engineering effort without AI assistance is adapted here to quantify output in the Codex study using an ensemble of models.
Codex: Codex is OpenAI’s public repository focused on AI-related code development. The analysis centers on pull requests from its 41 core contributors to measure estimated human effort and detect signs of AI-driven productivity gains.
OpenAI: OpenAI is an artificial intelligence research and deployment company that develops advanced models and maintains public code repositories for collaborative projects. The news uses data from OpenAI’s Codex repository to examine how AI tools may be accelerating engineering work among its core contributors.
GPT-5.5: GPT-5.5 is a state-of-the-art language model optimized for coding and technical evaluation. It is one of the three models used to generate median time estimates for pull requests in the Q2 2026 analysis.
Opus 4.8: Opus 4.8 is an advanced language model designed for complex reasoning and analysis tasks. It contributes to the ensemble of LLM judges estimating replication time for code changes in OpenAI’s Codex repository.
Gemini 3.1 Pro: Gemini 3.1 Pro is a high-performance language model from Google with strong capabilities in code and reasoning benchmarks. It participates in the multi-model ensemble providing effort estimates for the Codex contributor data.

`json
{
“AI Integration”: “AI coding assistants are now utilized in AI labs to streamline internal engineering processes.”,
“Industry Shift”: “Frontier AI companies report enhanced iteration speed when using advanced models for various tasks.”,
“Evaluation Methods”: “Organizations are enhancing LLM-based methods to assess productivity in software development.”
}
`