qwen

Qwen

Browse models from Qwen
Models · 25
Qwen-Audio-3.0-TTS-Plus

Qwen-Audio-3.0-TTS-Plus

qwen/qwen-audio-3.0-tts-plus
10.82Ktokens
0%Cache Hit Rate

Qwen-Audio-3.0-TTS-Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

Input Type
Output Type
Input$20/M tokens
Output-
Context-
Max Output-
28.48Mtokens
0%Cache Hit Rate

The Qwen3-VL-Embedding model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.

Input Type
Output Type
Input$0.1/M tokens
Output$0/M tokens
Context-
Max Output-
6.60Mtokens
0%Cache Hit Rate

Qwen3 Rerank is a new proprietary model of the Qwen model family. Specifically designed for reranking tasks, built on the Qwen3 foundation model. Leveraging Qwen3’s robust multilingual text understanding capabilities, the series achieves state-of-the-art performance across multiple benchmarks for text embedding and reranking tasks.

Input Type
Output Type
Input$0.1/M tokens
Output$0/M tokens
Context-
Max Output-
3.72Mtokens
0%Cache Hit Rate

The Qwen3-VL-Rerank model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.

Input Type
Output Type
Input$0.1/M tokens
Output$0/M tokens
Context-
Max Output-
9.79Btokens
87.77%Cache Hit Rate

Qwen3.6-Plus is Alibaba’s next-generation Qwen large language model released on April 2, 2026. Compared with version 3.5, Qwen 3.6 has made significant overall performance improvements and has exhibited remarkably strong agent-oriented programming capabilities.

Input Type
Output Type
Input$0.5-2/M tokens
Output$3-6/M tokens
Context-
Max Output-
2.63Btokens
86.6%Cache Hit Rate

2.4-trillion-parameter MoE flagship delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days. Handles hundreds of specialized tasks across legal, financial, design, and other professional domains, producing production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. In long-horizon tasks, plans autonomously, iterates through closed feedback loops, and continuously evolves.

Input Type
Output Type
Input$2/M tokens
Output$6/M tokens
Context-
Max Output-
1.21Btokens
85.6%Cache Hit Rate

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.

Input Type
Output Type
Input$0.03-0.2/M tokens
Output$0.13-0.8/M tokens
Context-
Max Output-
2.70Mtokens
-Cache Hit Rate

Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments.

Input Type
Output Type
Input$0.000035/seconds
Output-
Context-
Max Output-
16.36Btokens
73.46%Cache Hit Rate

Qwen3.7-Plus is a multimodal generation model in the Qwen3.7 family that supports text, image, and video inputs with text output. It is designed for agent-oriented workloads and strong performance in coding, office productivity, and long-horizon autonomous execution.

Input Type
Output Type
Input$0.4-1.2/M tokens
Output$1.6-4.8/M tokens
Context-
Max Output-
24.52Btokens
84.02%Cache Hit Rate
50% OFF

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous execution. The model offers notable gains in coding and agentic performance over prior Qwen generations and supports explicit prompt caching for efficient repeated context use.

Input Type
Output Type
Input$1.25/M tokens
Output$3.75/M tokens
Context-
Max Output-
197.87Mtokens
89.39%Cache Hit Rate
Sunset: 2026/09/08

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning traces across multi-turn conversations and supports structured output and function calling. Access is available exclusively through the Alibaba Cloud Model Studio and Qwen Studio APIs; no open weights are provided.

Input Type
Output Type
Input$1.3-2/M tokens
Output$7.8-12/M tokens
Context-
Max Output-
1.55Btokens
84.6%Cache Hit Rate

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in above 256K tokens. Prompt caching is supported, with both explicit cache read and cache creation pricing.

Input Type
Output Type
Input$0.25-1/M tokens
Output$1.5-4/M tokens
Context-
Max Output-
4.44Btokens
81.29%Cache Hit Rate

Qwen3.5 native vision-language Flash model series, based on a hybrid architecture design, integrates linear attention mechanisms with sparse Mixture-of-Experts (MoE) models, achieving higher inference efficiency. The model's performance has made a leap forward compared to the 3 series in both pure text and multimodal capabilities; it offers fast response times, balancing inference speed and performance.

Input Type
Output Type
Input$0.1/M tokens
Output$0.4/M tokens
Context-
Max Output-
5.19Btokens
48.98%Cache Hit Rate

Qwen3.5 Native Visual Language Series Plus model, based on a hybrid architecture design, integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher reasoning efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current top-tier cutting-edge models, with model performance showing leaps and bounds of progress compared to the 3 series in both pure text and multimodal aspects.

Input Type
Output Type
Input$0.4-0.5/M tokens
Output$2.4-3/M tokens
Context-
Max Output-
1.03Btokens
0.29%Cache Hit Rate
Sunset: 2026/09/08

Qwen3-Max-Thinking is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It delivers higher accuracy in math, coding, logic, and science tasks, follows complex instructions in Chinese and English more reliably, reduces hallucinations, and produces higher-quality responses for open-ended Q&A, writing, and conversation. The model supports over 100 languages with stronger translation and commonsense reasoning, and is optimized for retrieval-augmented generation (RAG) and tool calling, though it does not include a dedicated “thinking” mode.

Input Type
Output Type
Input$1.2-3/M tokens
Output$6-15/M tokens
Context-
Max Output-
175.12Mtokens
6.12%Cache Hit Rate

The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.

Input Type
Output Type
Input$0.2-0.6/M tokens
Output$1.6-4.8/M tokens
Context-
Max Output-
655.07Mtokens
74.56%Cache Hit Rate
Sunset: 2026/09/08

Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.

Input Type
Output Type
Input$1-6/M tokens
Output$5-60/M tokens
Context-
Max Output-
607.37Mtokens
27.7%Cache Hit Rate

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" (<think> blocks).

Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench.

Input Type
Output Type
Input$0.28/M tokens
Output$1.11/M tokens
Context-
Max Output-
195.41Mtokens
58.97%Cache Hit Rate

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, programming, and logical inference, and a "non-thinking" mode for general-purpose conversation. The model is fine-tuned for instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.

Input Type
Output Type
Input$0.14/M tokens
Output$1.4/M tokens
Context-
Max Output-
92.38Mtokens
32.14%Cache Hit Rate

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).

Pricing for the Alibaba endpoints varies by context length. Once a request is greater than 128k input tokens, the higher pricing is used.

Input Type
Output Type
Input$1.25/M tokens
Output$5.01/M tokens
Context-
Max Output-
Qwen: Qwen3 235B A22B Thinking 2507

Qwen: Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507
24.81Mtokens
10.94%Cache Hit Rate

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144 tokens of context. This "thinking-only" variant enhances structured logical reasoning, mathematics, science, and long-form generation, showing strong benchmark performance across AIME, SuperGPQA, LiveCodeBench, and MMLU-Redux. It enforces a special reasoning mode (</think>) and is designed for high-token outputs (up to 81,920 tokens) in challenging domains.

The model is instruction-tuned and excels at step-by-step reasoning, tool use, agentic workflows, and multilingual tasks. This release represents the most capable open-source variant in the Qwen3-235B series, surpassing many closed models in structured reasoning use cases.

Input Type
Output Type
Input$0.28/M tokens
Output$2.78/M tokens
Context-
Max Output-
3.41Mtokens
-Cache Hit Rate

Qwen-Image-2.0-Pro is an image generation model launched by Qwen AI.

Input Type
Output Type
Input-
Output$0.073/counts
Context-
Max Output-
3.00Mtokens
-Cache Hit Rate

Qwen-Image-2.0 is an image generation model launched by Qwen AI.

Input Type
Output Type
Input-
Output$0.0289/counts
Context-
Max Output-
96.00Ktokens
-Cache Hit Rate

Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool.

Input Type
Output Type
Input$0.003/counts
Output$0.04-0.075/counts
Context-
Max Output-
130.00Ktokens
-Cache Hit Rate

Accurate Prompt Understanding: Supports up to 4.5k token inputs, accurately interpreting complex text-and-image prompts and generating dense layouts in one pass. Reliable Text Rendering: Delivers crisp 10px text across 12 languages and 20+ fonts, making infographics and interfaces ready to use. Efficient Batch Production: Scales everyday tasks like posters, web pages, and UI screens with better cost efficiency, keeping creative output sustainable. Qwen-Image-3.0 Standard is built not just for visual quality, but for smooth, reliable daily creation—turning image generation into sustainable content productivity.

Input Type
Output Type
Input-
Output$0.03/counts
Context-
Max Output-