Alibaba's Qwen team on August 26, 2026 released Qwen3.8-Flash — a multimodal mixture-of-experts model billed as "ultimate cost efficiency" and a preview of the coming Qwen4 architecture.
125B total, 6B active, and an N-gram embedding trick
Qwen3.8-Flash carries 125 billion total parameters but activates only 6 billion per token. Another 51 billion parameters sit in a novel N-gram embedding layer that stores common phrases as standalone entries and can run in system RAM instead of GPU memory — one of the architecture innovations slated for Qwen4. The model natively supports a 262,144-token context window and can scale to one million tokens with YaRN. The open-weight variant, Qwen3.8-Flash-Next, is on Hugging Face and ModelScope; the production model ships through QwenCloud API at $0.16 per million input tokens and $0.47 per million output tokens.
Benchmark wins over DeepSeek and Claude
Alibaba's published benchmarks pit Flash against DeepSeek-V4-Flash and Anthropic's Claude Opus 4.6, both meaningfully larger or pricier. Flash led on agentic-coding tests DeepSWE (58.7) and SWE-bench Pro (62.5) and posted a wide gap on office/workflow benchmarks — 73.9 on CoWorkBench versus DeepSeek's 45.1, and 55.7 on JobBench, nearly double Qwen3.7-Plus. The Qwen team says the whole run cost roughly one-ninth of what Qwen3.7-Plus needed.
Pricing pressure keeps ratcheting up
Flash lists at about one-twelfth the input token price of the current Qwen3.8-Max flagship, following the earlier Qwen3.8-27B open-weight release and Z.AI's GLM-5.3 launch. The combined effect is more cost pressure on OpenAI and Anthropic — pressure OpenAI met last month with steep discounts on the GPT-5.6 line.
Reporting based on coverage from The Decoder, Bloomberg, and Qwen's own release notes.
