Alibaba has launched Qwen3.8-Max, a powerful AI model featuring 2.4 trillion parameters, designed to enhance performance by activating only a fraction of its capabilities per token. Utilizing a sparse mixture-of-experts architecture, this model can efficiently manage substantial computations, activating around 95 billion parameters while holding vast stored knowledge. Qwen3.8-Max has demonstrated impressive results in various benchmarks, including scoring 86.6 on Terminal Bench 2.1 and completing a 10-day autonomous coding task with impressive efficiency in software development, reflecting the growing trend of AI systems being tested on long-running, real-world applications. Additionally, it showcases advanced capabilities in interpreting aerial and satellite imagery, confirming its utility in specialized fields.

Alibaba: Alibaba is a major Chinese technology conglomerate with significant operations in cloud computing, e-commerce, and artificial intelligence research and development. The company has been actively advancing its Qwen series of large language models through its research teams. In this development, Alibaba announced the release of Qwen3.8-Max as part of its efforts to deliver high-performance AI systems with efficient inference capabilities.
Qwen3.8-Max: Qwen3.8-Max is Alibaba’s latest large-scale AI model built on a sparse mixture-of-experts architecture designed for efficient activation of parameters during inference. It supports extended context handling and excels in autonomous coding, multimodal tasks, and complex problem-solving benchmarks. The model was introduced in the recent announcement highlighting its competitive positioning in coding arenas and practical application scenarios such as image interpretation and web development.

Evaluation: Models are increasingly tested on autonomous long-running tasks that simulate real-world software development workflows.
Multimodal: Advanced AI systems are demonstrating strong capabilities in interpreting aerial and satellite imagery for specialized applications.
Architecture: Sparse mixture-of-experts designs allow AI models to scale stored knowledge while limiting active computation per token.