Alibaba Qwen Ships Qwen3.8-Omni-Flash With 1M-Token Multimodal Context

Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026 — a native omnimodal model with a 1M-token context, agentic video perception and audio input costs 98% below its predecessor.

Alibaba Qwen Ships Qwen3.8-Omni-Flash With 1M-Token Multimodal Context

Alibaba's Qwen team on September 18, 2026 released Qwen3.8-Omni-Flash, a native omnimodal model that ingests text, image, audio and video and returns text output at what the company describes as a "small enough for agents, large enough for real work" cost profile. The model launches API-only through QwenCloud, Alibaba Cloud Model Studio and Qwen Studio, and slots into Alibaba's growing Qwen family below the frontier Qwen3.8 tier we covered previously.

1M Context, 113 Languages, Two-Hour Video

Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture and ships with a 1M-token context window (991K max input, 131K output). It handles video files up to two hours or 2 GB via URL, audio clips up to three hours and audio understanding across 113 languages and dialects. Maximum reasoning length is set at 262K tokens with the highest xhigh thinking budget enabled by default.

Agentic multimodal AI representative image

Agentic Video Perception Cuts Tokens 45%

The most substantive architectural change is what Qwen calls agentic video perception: rather than watching a full clip end to end, the agent "starts from the question," then decides which segments to inspect and hear, iterating coarse-to-fine. Qwen reports a 45.7% token reduction for typical video workloads and an OmniVideoBench score climbing from 63.4 to 67.8 versus the same base model without agentic perception. On aggregate across 29 evaluations, Qwen3.8-Omni-Flash lands more than 25% above Qwen3.5-Omni-Plus, with a 36.5-point jump on WildClawBench-MM.

Pricing And Tooling

API pricing is set at $0.15 per million input tokens and $0.47 per million output tokens, with cache hits at $0.016 per million — a roughly 98% reduction in audio-input cost versus prior Qwen omni models. The launch bundles function calling, structured outputs, web search, context caching and batch processing, plus an Apache 2.0 Qwen-MM-Plugins pack that offers turnkey recipes such as video-to-notes and speaker-preserving translation. No open-weights or self-hosting option was announced, unlike Qwen's earlier open-weights Qwen3.8-27B release.

Why It Matters

Omnimodal, tool-calling models with million-token context are exactly the substrate that agent-heavy stacks need. Alibaba is betting that keeping this tier API-only — while Qwen3.8-Flash and the open-weight 27B carry the developer flywheel — lets it price aggressively against Google's Gemini 3.8 Flash and OpenAI's GPT-6 line in agentic audio-video workloads, particularly for enterprise buyers building call-center, meeting-notes and video-forensics pipelines.

Reporting based on coverage from MarkTechPost, TechNode, Pandaily and Qwen team materials.

Category: AI & Technology

Tags: AI AI Foundation Models AI Agents agentic AI

Related Articles