Databricks announced that its open-source coding model, GLM-5.2, has demonstrated the ability to compete with elite closed coding models, achieving top-tier performance on internal benchmarks and statistically tying in quality with Claude Opus 4.8. This evaluation was conducted using real internal pull requests and a multi-million-line codebase to create a more accurate testing environment, which revealed that GLM-5.2 operates at a cost of $1.28 per task, compared to Opus 4.8’s $1.94. The test highlighted how different harnesses can significantly impact coding agent costs by effectively managing the context provided to the model. As open models like GLM-5.2 continue to advance, they are now capable of tackling complex coding tasks, positioning them as viable alternatives in enterprise settings.

GLM-5.2: GLM-5.2 is an open-weight coding model designed for high-complexity software development tasks. In this news, it achieved statistically comparable quality to top closed models in Databricks’ private enterprise test, demonstrating that open models can now handle frontier-level coding work when paired with efficient harnesses.
Ali Ghodsi: Ali Ghodsi is the CEO of Databricks. In this news, he publicly shared the company’s findings from its internal model evaluation, emphasizing the value of testing on proprietary workloads and the cost savings achievable through different harnesses for the same models.
Databricks: Databricks is a cloud-based data and AI platform company focused on helping enterprises manage, analyze, and operationalize large-scale data and machine learning workloads. In this news, Databricks developed and ran a proprietary internal evaluation using its own multi-million-line enterprise codebase and real engineering tasks to compare coding models. The company highlighted GLM-5.2’s strong performance and the role of harnesses in controlling costs.
Claude Opus 4.8: Claude Opus 4.8 is a leading closed-source coding model developed by Anthropic for advanced reasoning and software tasks. In this news, it served as a key benchmark in Databricks’ evaluation where GLM-5.2 tied with it on quality for real internal code changes and tests.

Harness Efficiency: The choice of harness can substantially affect coding agent costs by reducing repeated context sent to the model while maintaining output quality.
Internal Benchmarking: Databricks created a private evaluation harness using real PRs, tests, and its own enterprise codebase to assess models more reliably than public benchmarks.
Open Model Capability: Open models like GLM-5.2 have reached a point where they can manage the highest difficulty coding tasks in realistic enterprise environments.