Cerebras Accelerates AI Inference with Alibaba's Qwen 3.8 27B Model
Cerebras now serves Alibaba's open-weight Qwen 3.8 27B model on its public inference endpoints, achieving approximately 1,500 tokens per second. This integration, which became available around September 3-4, 2026, offers a significant speed advantage for a 27-billion-parameter model, surpassing typical GPU-based API serving.
High-Performance Inference for Qwen 3.8 27B
The Qwen 3.8 27B model, a 27-billion-parameter multimodal AI, is now operational on Cerebras' inference platform. This model is designed for image and video understanding, alongside controllable reasoning capabilities. Its performance on Cerebras' infrastructure is particularly noteworthy, achieving speeds of approximately 1,500 tokens per second. This throughput is described as roughly an order of magnitude faster than what is typically observed with GPU-based API serving for models of this scale.
Cerebras emphasizes its approach to serving original, unpruned models, utilizing selective weight-only storage quantization to maintain model integrity while optimizing for speed. This strategy aims to deliver high-fidelity AI inference without compromising on the model's inherent capabilities.
Context Window and Pricing Details
The Qwen 3.8 27B model offers varying context window sizes depending on the user tier. A 64k-token context window is available for users on the free trial tier, while paid tiers provide an extended 128k-token context window. This flexibility allows users to process longer sequences of data, which is crucial for complex AI applications.
Pricing for the Qwen 3.8 27B API on Cerebras is set at $0.99 per million input tokens and $1.49 per million output tokens. This structure provides a clear cost model for developers integrating the service into their applications.
Comparison with Other Cerebras Offerings
Cerebras also hosts the OpenAI GPT OSS 120B model on its public endpoints, which operates at approximately 3,000 tokens per second. This model features context windows of 65k and 131k tokens. The performance of Qwen 3.8 27B, while impressive, positions it as a strong contender in Cerebras' portfolio, especially given its multimodal capabilities and efficient parameter count.
The launch of Qwen 3.8 27B on Cerebras' platform garnered significant attention within the developer community, including substantial discussion on Hacker News, indicating strong interest in high-performance, open-weight AI models.
Implications for AI Development
The integration of Alibaba's Qwen 3.8 27B on Cerebras' high-speed inference platform offers developers a powerful new option for deploying multimodal AI applications. The reported speed advantages could enable more responsive and efficient AI-powered services, particularly for tasks requiring rapid processing of large datasets or complex reasoning. This development underscores the ongoing advancements in AI infrastructure aimed at making sophisticated models more accessible and performant.
Conclusion
Cerebras' addition of Alibaba's Qwen 3.8 27B model to its public inference endpoints marks a significant step in providing high-speed, multimodal AI capabilities. With its impressive token throughput and flexible context windows, the Qwen 3.8 27B model on Cerebras offers a compelling solution for developers seeking efficient and powerful AI inference. The continued expansion of such offerings on platforms like Cerebras will likely drive further innovation in the application of advanced AI models.
Sources
- Qwen3.8-27B Cerebras API Pricing: $0.99/$1.49 | AI Pricing Guru
- Model Catalog - Cerebras Inference
- Cerebras Serves Qwen 3.8 27B at 1,500 Tokens/s — AICrier
- 💡 Exciting news! Cerebras will begin hosting Qwen 3.8...
- Cerebras on X: "Congratulations to @Alibaba_Qwen team on the Qwen 3.8 27B launch! We can't wait for our users to try it out! Contact us for dedicated deployments and it's going to be up soon in the Cerebras Shared Tier. 🟧⚡️" / X
Recommended AI tools
Google Cloud Vertex AI
Data Analytics
Gemini, Vertex AI, and AI infrastructure—everything you need to build and scale enterprise AI on Google Cloud.
Google AI Studio
Productivity & Collaboration
The fastest way to build AI-first applications with Google Gemini.
Hugging Face
Scientific Research
Democratizing good machine learning, one commit at a time.
Transformers
Conversational AI
State-of-the-art AI models for text, vision, audio, video & multimodal—open-source tools for everyone.
Venice AI
Conversational AI
Private AI for Unlimited Creative Freedom
fal.ai
Image Generation
Empowering AI for Everyone
About the Author

Albert Schaper is a co-founder of Best-AI.org. He focuses on product strategy, AI adoption, practical tool selection, and educational content that helps users compare AI products with clearer context.
More from AlbertWas this article helpful?
Found outdated info or have suggestions? Send us a note.