Qwen API Pricing 2026: The Complete Alibaba Cloud Model Studio Cost Guide
๐ Data note (Sep 1, 2026): All rates below are synced from the official Model Studio model inference pricing page (last updated Aug 31, 2026) and the Chinese ็พ็ผๆจกๅไปทๆ ผ้กต. Standard list prices are shown; limited-time promotio
๐ Data note (Sep 1, 2026): All rates below are synced from the official Model Studio model inference pricing page (last updated Aug 31, 2026) and the Chinese ็พ็ผๆจกๅไปทๆ ผ้กต. Standard list prices are shown; limited-time promotions (such as the Qwen3.7-Max 50% discount) are flagged. Always confirm final numbers in the console before production budgeting.
1. How Model Studio Billing Works
Alibaba Cloud Model Studio (the international name for Bailian / ็พ็ผ) is Alibabaโs one-stop LLM platform: Qwen text models, DeepSeek, Qwen-VL vision models, QwQ reasoning models, embedding, rerank, speech and image generation, all behind an OpenAI-compatible API. Billing is straightforward:
Pay-as-you-go by tokens. Fee = input tokens ร input price + output tokens ร output price. Prices are quoted per 1 million tokens.
Tiered pricing on some models. For models with context tiers, the price is set by the total input tokens of that single request โ and all tokens in the request are billed at that tierโs rate. A 100K-input request on a two-tier model (0โ32K, 32Kโ128K) is billed entirely at the 32Kโ128K rate.
Batch inference = 50% off. Models supporting batch calls charge half price for both input and output, with results returned asynchronously.
Context caching = up to ~90% off cached input. Explicit cache creation costs 125% of the input price, but cache-hit input tokens cost about 10%. Repeated system prompts or long reference documents become dramatically cheaper.
Off-peak / night discounts. Some models carry automatic night discounts (22:00โ08:00 Beijing time, UTC+8) with no signup โ currently flagged on Qwen3.7 series list prices.
Free quota. New accounts get 1 million tokens per model (input and output each), valid for 90 days from activation โ across dozens of models this totals roughly 70 million free tokens. On the international (Singapore) deployment, the free quota applies to the models listed there; mainland deployment (Beijing) has its own free-quota list.
2. Qwen API Price Table โ International Deployment (USD, per 1M tokens)
These are the International scope rates (Singapore region) most overseas developers use. Standard real-time prices; batch halves them where supported.
| Model | Input (USD/1M) | Output (USD/1M) | Context | Best for |
|---|---|---|---|---|
| Qwen3.8-Max (flagship, GA Aug 3, 2026) | $2.00 flat | $6.00 flat | 1M | Hardest agentic/coding tasks; one flat rate at any length |
| Qwen3.7-Max | $2.50 list (50% off โ $1.25) | $7.50 list (50% off โ $3.75) | 1M | Previous flagship โ cheapest high-end while promo lasts |
| Qwen3-Max | $1.20 โ $2.40 โ $3.00 | $6 โ $12 โ $15 | 262K | Reasoning & coding (tiers at 32K / 128K input) |
| Qwen3.7-Plus | $0.48 (โค256K: $1.44) | $1.92 (โค256K: $5.76) | 1M | Multimodal mid-tier, production workhorse |
| Qwen-Plus (qwen-plus-2025-12-01) | from ~$0.40 | from ~$1.20 | 1M | The default โstart hereโ model for most apps |
| Qwen-Flash | from $0.05 | from $0.40 | 1M | High-volume classification, tagging, extraction |
| Qwen3.7-Flash | ~$0.03โ$0.07 | ~$0.13โ$0.20 | 1M | Ultra-cheap batch processing |
| Qwen-Turbo (legacy) | ~$0.05 | ~$0.20 | 1M | No longer updated โ use Qwen-Flash for new projects |
| QwQ-Plus (reasoning) | ~$0.82 (ยฅ5.871 intl) | ~$2.47 (ยฅ17.614 intl) | 128K | Deep thinking mode; free 1M tokens each |
| Qwen-Long (long context) | ~$0.07 (ยฅ0.5 mainland) | ~$0.28 (ยฅ2 mainland) | 1M | Whole-document analysis on a budget |
Source: Alibaba Cloud Model Studio โ Model inference pricing (Aug 31, 2026). CNY-converted rows marked with ยฅ use the mainland rate at ~7.2 CNY/USD as an approximation; the international rate is billed in USD.
3. Qwen API Price Table โ Mainland China Deployment (CNY, per 1M tokens)
For China-facing products served from Beijing (North China 2). These rates include the models most commonly used by Chinese teams:
| ๆจกๅ Model | ่พๅ ฅ Input (ยฅ/1M) | ่พๅบ Output (ยฅ/1M) | ๅ ่ดน้ขๅบฆ Free quota |
|---|---|---|---|
| qwen-turbo / qwen-turbo-latest | ยฅ0.367 | ยฅ1.468 | Batch ๅไปท |
| qwen-plus-2025-04-28 ๅๆฉๆ็ | ยฅ0.8 | ยฅ2 | ๅ 100 ไธ tokens |
| qwen3.6-plus (โค32K ๆกฃ) | ยฅ2 | ยฅ12 | ๅ 100 ไธ tokens |
| qwq-plus (ๆ่ๆจกๅผ) | ยฅ1.6 | ยฅ4 | ๅ 100 ไธ tokens |
| qwen-long | ยฅ0.5 | ยฅ2 | ๅ 100 ไธ tokens |
| qvq-plus (่ง่งๆจ็) | ยฅ2 | ยฅ5 | ๅ 100 ไธ tokens |
| qvq-max | ยฅ8 | ยฅ32 | ๅ 100 ไธ tokens |
| qwen3-vl-flash (โค32K) | ยฅ0.15 | ยฅ1.5 | Batch ๅไปท |
Source: ้ฟ้ไบๅธฎๅฉไธญๅฟ โ ็พ็ผๆจกๅไปทๆ ผ. Mainland free quota: 1 million tokens each for input/output, valid 90 days after Bailian activation.
4. How to Pick a Model (and What It Costs You)
High-volume simple work (classification, tagging, extraction, chat triage): Qwen-Flash. At $0.05/$0.40 per million, 10 million input + 2 million output tokens cost roughly $1.30/day โ under $40/month for a busy bot.
Mainstream app workhorse (summarization, drafting, RAG answers, coding assist): Qwen-Plus / Qwen3.7-Plus. 10M in + 2M out on Qwen3.7-Plus works out to about $8.6/month at standard rates โ this is why Plus is the default for production.
Agentic and hard reasoning (multi-step tools, codebase agents, math): Qwen3.8-Max. Flat $2/$6 across the full 1M context means no long-prompt cliff; the same 10M+2M workload is about $32/month.
Batch/night jobs: batch API halves everything; night discounts on Qwen3.7 can reach 80% off list. Offline pipelines should never run at peak real-time price.
DeepSeek is available on the same platform โ DeepSeek-V4-Flash and V4-Pro are callable through Model Studio too, so you can A/B vendors without changing infrastructure.
5. Worked Examples: Real Monthly Bills
| Scenario | Volume (monthly) | Model | Est. bill |
|---|---|---|---|
| Support chatbot, 50K conversations | 20M input / 4M output | Qwen-Flash | ~$2.6/month |
| RAG knowledge assistant, 100K queries | 50M input / 10M output | Qwen3.7-Plus | ~$43/month |
| Coding agent, 5K heavy tasks | 30M input / 10M output | Qwen3.8-Max | ~$120/month |
| Same coding agent, batch mode | 30M input / 10M output | Qwen3.8-Max batch | ~$60/month |
| Document processing, long-context | 100M input / 5M output | Qwen-Long (mainland ยฅ) | ~ยฅ60 (~$8)/month |
The takeaway: outside of heavy agent workloads, most production apps run on tens of dollars a month โ and the free quota covers the first ~70 million tokens of experimentation entirely.
6. Free Tokens: What New Accounts Actually Get
1 million tokens per model, both input and output, for each model in the free-quota list.
Valid 90 days from Model Studio activation (or model release / application approval, whichever is later).
International deployment: free quota is granted in the Singapore region; other international regions donโt carry it.
Mainland deployment: free quota in the Beijing (North China 2) region.
The old unlimited free developer tier ended April 15, 2026 โ the current program is this per-model trial pack.
7. Three Ways to Cut the Bill Further
Prompt caching. If your system prompt + retrieved context repeats across users, explicit cache hits drop input cost to ~10%. RAG apps routinely cut 60โ80% of input spend this way.
Batch calls for non-interactive work. Evals, embeddings-style sweeps, nightly summaries: 50% off with no quality difference.
Model routing. Use Flash for triage and simple intents, escalate only the hard 5โ10% of requests to Plus or Max. Most โMax-onlyโ apps waste 80%+ of their budget.
8. Getting Started in 10 Minutes
Register an Alibaba Cloud international account and claim the $200 starter credit (new users; approval ~3 business days where required).
Open the Model Studio console, activate the service โ free tokens are granted automatically.
Create an API key and call the OpenAI-compatible endpoint (https://dashscope-intl.aliyuncs.com/compatible-mode/v1) with your existing OpenAI SDK โ just change base URL and key.
Start on Qwen-Flash or Qwen-Plus, switch to Qwen3.8-Max only where quality demands it, and turn on caching once prompts stabilize.
FAQ
Is Qwen API free?
The hosted API is pay-per-token, but new accounts get ~1 million free tokens per model (roughly 70 million total) for 90 days. Most Qwen models are also open-weight, so you can self-host on GPU servers at zero per-token cost for steady high volume.
How much does Qwen3.8-Max cost?
$2.00 per million input tokens and $6.00 per million output tokens on international deployment, flat across the entire 1M-token context window. Batch calls halve this to $1/$3-ish effective rates.
Qwen vs DeepSeek โ which is cheaper?
DeepSeek-V4-Flash undercuts Qwen-Flash on output price; Qwen-Plus is slightly cheaper on input. Both are callable from the same Model Studio account, so run your own A/B โ price differences are smaller than quality-fit differences for most workloads.
Do prices change often?
Alibaba cuts prices aggressively as new models ship โ Qwen3.7-Max is currently 50% off list and newer Flash tiers keep dropping. Budget against list prices for long-term planning and treat promotions as upside.
Start Building with Qwen Today
$200 international credit ยท ~70 million free Qwen tokens ยท OpenAI-compatible API ยท Singapore + Beijing regions
Claim $200 Free Credit
Open Model Studio
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.