The latest open-source model, GLM-5.2 from @Zai_org, has been deployed in Agent Arena, a platform designed for evaluating AI agents through live user sessions on real tasks. In Agent Arena, models like GLM-5.2 utilize a variety of tools, including web search and filesystem access, to perform complex workflows such as writing code and analyzing documents. The performance of models is tracked on a leaderboard that evaluates their effectiveness across multiple metrics, including task success and user satisfaction, based on data from over 300,000 tasks and 40 million lines of code.

Zai: Zai is an AI organization focused on developing advanced large language models, including the GLM series. Its latest release, GLM-5.2, is highlighted in the news for participating in Agent Arena evaluations. The model is being tested on complex agentic tasks alongside offerings from leading AI labs.
OpenAI: OpenAI develops frontier AI models and systems designed for a wide range of applications. In the news, its GPT-5.5 model ranks at the top of the Agent Arena leaderboard for agentic performance. This positioning underscores its role in advancing real-world workflow automation benchmarks.
Anthropic: Anthropic creates AI models with an emphasis on safety and advanced reasoning capabilities. The news places its Claude-Opus-4.7 model in second position on the Agent Arena leaderboard. This reflects its contributions to agent evaluations involving tool use and iterative task completion.
Kimi Moonshot: Kimi Moonshot develops large language models tailored for diverse practical uses. The news lists its Kimi-K2.6 model among the top performers in Agent Arena. This demonstrates its involvement in scaling evaluations of models handling complex, tool-augmented workflows.
Google DeepMind: Google DeepMind conducts research and development on cutting-edge AI technologies and multimodal systems. Its Gemini-3.1-Pro model appears in the Agent Arena rankings according to the provided leaderboard snapshot. The inclusion highlights ongoing competition in agentic AI assessment environments.

Arena Variants: Models like GLM-5.2 can also be tested in related environments such as Text Arena and Code Arena to assess performance on specialized real-world tasks.
Agent Evaluation: Agent Arena measures model performance through live user sessions that involve real tasks and tool interactions such as web search, filesystem access, and terminal commands.
Leaderboard Signals: Rankings in Agent Arena incorporate multiple signals including task success, steerability, error recovery, user feedback, and tool hallucination rates.