Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 6 min read

Qwen API Pricing 2026: The Complete Alibaba Cloud Model Studio Cost Guide

๐Ÿ“… Data note (Sep 1, 2026): All rates below are synced from the official Model Studio model inference pricing page (last updated Aug 31, 2026) and the Chinese ็™พ็‚ผๆจกๅž‹ไปทๆ ผ้กต. Standard list prices are shown; limited-time promotio

๐Ÿ“… Data note (Sep 1, 2026): All rates below are synced from the official Model Studio model inference pricing page (last updated Aug 31, 2026) and the Chinese ็™พ็‚ผๆจกๅž‹ไปทๆ ผ้กต. Standard list prices are shown; limited-time promotions (such as the Qwen3.7-Max 50% discount) are flagged. Always confirm final numbers in the console before production budgeting.

1. How Model Studio Billing Works

Alibaba Cloud Model Studio (the international name for Bailian / ็™พ็‚ผ) is Alibabaโ€™s one-stop LLM platform: Qwen text models, DeepSeek, Qwen-VL vision models, QwQ reasoning models, embedding, rerank, speech and image generation, all behind an OpenAI-compatible API. Billing is straightforward:

  • Pay-as-you-go by tokens. Fee = input tokens ร— input price + output tokens ร— output price. Prices are quoted per 1 million tokens.

  • Tiered pricing on some models. For models with context tiers, the price is set by the total input tokens of that single request โ€” and all tokens in the request are billed at that tierโ€™s rate. A 100K-input request on a two-tier model (0โ€“32K, 32Kโ€“128K) is billed entirely at the 32Kโ€“128K rate.

  • Batch inference = 50% off. Models supporting batch calls charge half price for both input and output, with results returned asynchronously.

  • Context caching = up to ~90% off cached input. Explicit cache creation costs 125% of the input price, but cache-hit input tokens cost about 10%. Repeated system prompts or long reference documents become dramatically cheaper.

  • Off-peak / night discounts. Some models carry automatic night discounts (22:00โ€“08:00 Beijing time, UTC+8) with no signup โ€” currently flagged on Qwen3.7 series list prices.

  • Free quota. New accounts get 1 million tokens per model (input and output each), valid for 90 days from activation โ€” across dozens of models this totals roughly 70 million free tokens. On the international (Singapore) deployment, the free quota applies to the models listed there; mainland deployment (Beijing) has its own free-quota list.

2. Qwen API Price Table โ€” International Deployment (USD, per 1M tokens)

These are the International scope rates (Singapore region) most overseas developers use. Standard real-time prices; batch halves them where supported.

Model Input (USD/1M) Output (USD/1M) Context Best for
Qwen3.8-Max (flagship, GA Aug 3, 2026) $2.00 flat $6.00 flat 1M Hardest agentic/coding tasks; one flat rate at any length
Qwen3.7-Max $2.50 list (50% off โ†’ $1.25) $7.50 list (50% off โ†’ $3.75) 1M Previous flagship โ€” cheapest high-end while promo lasts
Qwen3-Max $1.20 โ†’ $2.40 โ†’ $3.00 $6 โ†’ $12 โ†’ $15 262K Reasoning & coding (tiers at 32K / 128K input)
Qwen3.7-Plus $0.48 (โ‰ค256K: $1.44) $1.92 (โ‰ค256K: $5.76) 1M Multimodal mid-tier, production workhorse
Qwen-Plus (qwen-plus-2025-12-01) from ~$0.40 from ~$1.20 1M The default โ€œstart hereโ€ model for most apps
Qwen-Flash from $0.05 from $0.40 1M High-volume classification, tagging, extraction
Qwen3.7-Flash ~$0.03โ€“$0.07 ~$0.13โ€“$0.20 1M Ultra-cheap batch processing
Qwen-Turbo (legacy) ~$0.05 ~$0.20 1M No longer updated โ€” use Qwen-Flash for new projects
QwQ-Plus (reasoning) ~$0.82 (ยฅ5.871 intl) ~$2.47 (ยฅ17.614 intl) 128K Deep thinking mode; free 1M tokens each
Qwen-Long (long context) ~$0.07 (ยฅ0.5 mainland) ~$0.28 (ยฅ2 mainland) 1M Whole-document analysis on a budget

Source: Alibaba Cloud Model Studio โ€” Model inference pricing (Aug 31, 2026). CNY-converted rows marked with ยฅ use the mainland rate at ~7.2 CNY/USD as an approximation; the international rate is billed in USD.

3. Qwen API Price Table โ€” Mainland China Deployment (CNY, per 1M tokens)

For China-facing products served from Beijing (North China 2). These rates include the models most commonly used by Chinese teams:

ๆจกๅž‹ Model ่พ“ๅ…ฅ Input (ยฅ/1M) ่พ“ๅ‡บ Output (ยฅ/1M) ๅ…่ดน้ขๅบฆ Free quota
qwen-turbo / qwen-turbo-latest ยฅ0.367 ยฅ1.468 Batch ๅŠไปท
qwen-plus-2025-04-28 ๅŠๆ—ฉๆœŸ็‰ˆ ยฅ0.8 ยฅ2 ๅ„ 100 ไธ‡ tokens
qwen3.6-plus (โ‰ค32K ๆกฃ) ยฅ2 ยฅ12 ๅ„ 100 ไธ‡ tokens
qwq-plus (ๆ€่€ƒๆจกๅผ) ยฅ1.6 ยฅ4 ๅ„ 100 ไธ‡ tokens
qwen-long ยฅ0.5 ยฅ2 ๅ„ 100 ไธ‡ tokens
qvq-plus (่ง†่ง‰ๆŽจ็†) ยฅ2 ยฅ5 ๅ„ 100 ไธ‡ tokens
qvq-max ยฅ8 ยฅ32 ๅ„ 100 ไธ‡ tokens
qwen3-vl-flash (โ‰ค32K) ยฅ0.15 ยฅ1.5 Batch ๅŠไปท

Source: ้˜ฟ้‡Œไบ‘ๅธฎๅŠฉไธญๅฟƒ โ€” ็™พ็‚ผๆจกๅž‹ไปทๆ ผ. Mainland free quota: 1 million tokens each for input/output, valid 90 days after Bailian activation.

4. How to Pick a Model (and What It Costs You)

  • High-volume simple work (classification, tagging, extraction, chat triage): Qwen-Flash. At $0.05/$0.40 per million, 10 million input + 2 million output tokens cost roughly $1.30/day โ€” under $40/month for a busy bot.

  • Mainstream app workhorse (summarization, drafting, RAG answers, coding assist): Qwen-Plus / Qwen3.7-Plus. 10M in + 2M out on Qwen3.7-Plus works out to about $8.6/month at standard rates โ€” this is why Plus is the default for production.

  • Agentic and hard reasoning (multi-step tools, codebase agents, math): Qwen3.8-Max. Flat $2/$6 across the full 1M context means no long-prompt cliff; the same 10M+2M workload is about $32/month.

  • Batch/night jobs: batch API halves everything; night discounts on Qwen3.7 can reach 80% off list. Offline pipelines should never run at peak real-time price.

  • DeepSeek is available on the same platform โ€” DeepSeek-V4-Flash and V4-Pro are callable through Model Studio too, so you can A/B vendors without changing infrastructure.

5. Worked Examples: Real Monthly Bills

Scenario Volume (monthly) Model Est. bill
Support chatbot, 50K conversations 20M input / 4M output Qwen-Flash ~$2.6/month
RAG knowledge assistant, 100K queries 50M input / 10M output Qwen3.7-Plus ~$43/month
Coding agent, 5K heavy tasks 30M input / 10M output Qwen3.8-Max ~$120/month
Same coding agent, batch mode 30M input / 10M output Qwen3.8-Max batch ~$60/month
Document processing, long-context 100M input / 5M output Qwen-Long (mainland ยฅ) ~ยฅ60 (~$8)/month

The takeaway: outside of heavy agent workloads, most production apps run on tens of dollars a month โ€” and the free quota covers the first ~70 million tokens of experimentation entirely.

6. Free Tokens: What New Accounts Actually Get

  • 1 million tokens per model, both input and output, for each model in the free-quota list.

  • Valid 90 days from Model Studio activation (or model release / application approval, whichever is later).

  • International deployment: free quota is granted in the Singapore region; other international regions donโ€™t carry it.

  • Mainland deployment: free quota in the Beijing (North China 2) region.

  • The old unlimited free developer tier ended April 15, 2026 โ€” the current program is this per-model trial pack.

7. Three Ways to Cut the Bill Further

  • Prompt caching. If your system prompt + retrieved context repeats across users, explicit cache hits drop input cost to ~10%. RAG apps routinely cut 60โ€“80% of input spend this way.

  • Batch calls for non-interactive work. Evals, embeddings-style sweeps, nightly summaries: 50% off with no quality difference.

  • Model routing. Use Flash for triage and simple intents, escalate only the hard 5โ€“10% of requests to Plus or Max. Most โ€œMax-onlyโ€ apps waste 80%+ of their budget.

8. Getting Started in 10 Minutes

  • Register an Alibaba Cloud international account and claim the $200 starter credit (new users; approval ~3 business days where required).

  • Open the Model Studio console, activate the service โ€” free tokens are granted automatically.

  • Create an API key and call the OpenAI-compatible endpoint (https://dashscope-intl.aliyuncs.com/compatible-mode/v1) with your existing OpenAI SDK โ€” just change base URL and key.

  • Start on Qwen-Flash or Qwen-Plus, switch to Qwen3.8-Max only where quality demands it, and turn on caching once prompts stabilize.

FAQ

Is Qwen API free?

The hosted API is pay-per-token, but new accounts get ~1 million free tokens per model (roughly 70 million total) for 90 days. Most Qwen models are also open-weight, so you can self-host on GPU servers at zero per-token cost for steady high volume.

How much does Qwen3.8-Max cost?

$2.00 per million input tokens and $6.00 per million output tokens on international deployment, flat across the entire 1M-token context window. Batch calls halve this to $1/$3-ish effective rates.

Qwen vs DeepSeek โ€” which is cheaper?

DeepSeek-V4-Flash undercuts Qwen-Flash on output price; Qwen-Plus is slightly cheaper on input. Both are callable from the same Model Studio account, so run your own A/B โ€” price differences are smaller than quality-fit differences for most workloads.

Do prices change often?

Alibaba cuts prices aggressively as new models ship โ€” Qwen3.7-Max is currently 50% off list and newer Flash tiers keep dropping. Budget against list prices for long-term planning and treat promotions as upside.

Start Building with Qwen Today

$200 international credit ยท ~70 million free Qwen tokens ยท OpenAI-compatible API ยท Singapore + Beijing regions

Claim $200 Free Credit
Open Model Studio

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.