Alibaba's Qwen3.8-Omni-Flash: 1M-Token Omnimodal AI with 98% Cheaper Audio API
Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, an omnimodal AI model with a 1M-token context window designed for audio-visual agents. This new model processes text, image, audio, and video inputs, aiming to enhance AI agents in planning tasks, tool calling, and creative work, while also reducing audio API costs by over 98%. For broader context, explore our AI Tools Pricing.
Introducing Qwen3.8-Omni-Flash: Omnimodal Capabilities
Qwen3.8-Omni-Flash is positioned as a native omnimodal model, available through the Qianwen AI platform and Alibaba Cloud Model Studio APIs. Its core design supports a wide range of input modalities, allowing for complex interactions that integrate various forms of data. The model's 1M-token context window facilitates processing extensive information, which is crucial for sophisticated AI agent applications.
A companion Realtime variant of Qwen3.8-Omni-Flash has also been introduced. This variant is specifically engineered to handle low-latency live audio-visual streams, including advanced spatial audio perception, catering to applications requiring immediate processing of dynamic, real-world data.
Significant API Pricing Reductions
Alibaba has substantially reduced API pricing for Qwen3.8-Omni-Flash. The cost for audio input has been cut by more than 98% per hour, while audio-visual input pricing has seen a reduction of over 93%. These price adjustments aim to make omnimodal AI capabilities more accessible for developers and businesses utilizing the Alibaba Cloud Model Studio APIs.
Performance Improvements and Agentic Design
Qwen3.8-Omni-Flash demonstrates notable performance enhancements compared to its predecessor, Qwen3.5-Omni-Plus. Across 29 evaluations, the new model achieved an average score improvement of more than 25%. This advancement is partly attributed to its agentic design, which optimizes token consumption for long-video tasks. This design reportedly reduces token usage by approximately 45.7% while simultaneously improving accuracy.
The model's speech recognition capabilities are extensive, covering 74 languages and 39 Chinese dialects. In a 12-hour autonomous training loop, Qwen3.8-Omni-Flash reduced the Sichuan-dialect speech recognition error rate of Qwen2.5-Omni-3B from 25.79% to 15.30%, indicating a significant improvement in handling specific linguistic variations.
Expanded Ecosystem: Qwen-Live Harness and Qwen-MM-Plugins
To support the deployment and interaction with Qwen3.8-Omni-Flash, Alibaba has open-sourced Qwen-Live Harness. This runtime environment is designed for continuous real-time omnimodal interaction, enabling developers to build more dynamic and responsive AI applications. The Qwen-MM-Plugins suite has also been expanded with two new tools: Video2Note and Omni Skill Creator. These additions aim to provide developers with more versatile options for integrating multimodal functionalities into their projects.
Conclusion
The release of Alibaba's Qwen3.8-Omni-Flash on September 18, 2026, marks an advancement in omnimodal AI, offering enhanced capabilities for processing diverse data types and significant reductions in API costs. With its improved performance, agentic design, and expanded ecosystem tools like Qwen-Live Harness and Qwen-MM-Plugins, the model is positioned to support the development of more sophisticated and efficient AI agents. Developers and organizations interested in leveraging advanced multimodal AI for planning, tool calling, and creative applications should consider exploring the new offerings on the Qianwen AI Platform and through Alibaba Cloud Model Studio APIs.
Sources
About the Author

Albert Schaper is a co-founder of Best-AI.org. He focuses on product strategy, AI adoption, practical tool selection, and educational content that helps users compare AI products with clearer context.
More from AlbertWas this article helpful?
Found outdated info or have suggestions? Send us a note.