Dev.to WebDev 🛠 Dev 👁 0 📖 3 min read

LLM Pricing Digest: Grok-4 Gets a Price Drop, and a Wave of New Models From Qwen, Xiaomi, and Z-AI

LLM Pricing Digest — Week of September 29, 2026 Weekly data from LLM Price Watch, which tracks live pricing for Claude, GPT, Gemini, DeepSeek, Grok, and more. The One Real Price Change: Grok-4 This week

LLM Pricing Digest — Week of September 29, 2026

Weekly data from LLM Price Watch, which tracks live pricing for Claude, GPT, Gemini, DeepSeek, Grok, and more.

The One Real Price Change: Grok-4

This week had one confirmed price movement: grok-4 dropped to $2/million input tokens and $6/million output tokens.

To put that in context, Grok-4 launched as a frontier-tier model, and pricing at that level used to reliably mean $10–$15+ per million on input. At $2 in and $6 out, it's now sitting in territory that makes it genuinely worth comparing against mid-tier options for production workloads — not just for experimentation.

Practically speaking: if you've been routing complex reasoning tasks to GPT-4o or Claude Sonnet and paying ~$5–$15/M input depending on your mix, Grok-4 at $2 input is worth a benchmark run. Output-heavy workloads (long generations, document drafting) will feel the $6/M output price more, but that's still competitive for a model at this capability tier.

No price changes detected for Claude, GPT, or Gemini models this week.

New Models: Qwen, Xiaomi, and Z-AI All Showed Up at Once

Fifteen new models appeared in our tracker on September 30th. They all landed the same day, which usually means a coordinated release push rather than a gradual rollout. Here's the breakdown:

Qwen (5 models)

  • qwen/qwen3.8-max-prime
  • qwen/qwen3.8-omni-flash
  • qwen/qwen3.8-max-0902
  • qwen/qwen3.8-flash
  • qwen/qwen3.8-27b
  • qwen/qwen3.8-27b:free ← free tier variant

Qwen continues to ship variants aggressively. The naming here follows their established pattern: max = higher capability, flash = faster/cheaper, omni = multimodal. The 27B free tier is notable — free access to a 27B model is useful for prototyping or low-volume use cases where cost is the primary constraint.

Xiaomi (3 models)

  • xiaomi/mimo-v2.6-pro-ultraspeed
  • xiaomi/mimo-v2.6-flash
  • xiaomi/mimo-v2.6-pro

Xiaomi's MiMo line is newer to the router ecosystem. Three variants in one drop — ultraspeed, flash, and pro — suggests they're covering the same speed/quality tradeoff spectrum everyone else is. No pricing data to share yet beyond what's in the tracker.

Z-AI / GLM-5.3 (7 models)

  • z-ai/glm-5.3-prime
  • z-ai/glm-5.3-flashx
  • z-ai/glm-5.3-flash
  • z-ai/glm-5.3-flash:batch
  • z-ai/glm-5.3
  • z-ai/glm-5.3:batch

Z-AI's GLM series (from Zhipu AI) is the most prolific of the three this week, with seven variants including batch-mode endpoints. Batch variants matter if you're running offline pipelines — they're typically priced lower in exchange for slower turnaround. Worth checking if you have async workloads that don't need real-time responses.

What This Week Means Practically

If you're actively managing API costs right now:

  1. Re-evaluate Grok-4 if you dismissed it earlier on price. $2 input is a different conversation than where it started.
  2. The Qwen 27B free tier is a legitimate option for dev/test environments or low-stakes internal tools.
  3. GLM-5.3 batch endpoints are worth a look if you're doing bulk processing — batch pricing tends to undercut synchronous endpoints meaningfully.
  4. Xiaomi MiMo is one to watch but probably wait for more community benchmarks before routing production traffic there.

The broader trend continues to hold: the number of capable models available through routing APIs keeps growing, and the middle of the market (the $1–$6/M input range) is getting crowded. That's good for developers with flexibility in model choice.

Full pricing data, historical charts, and alerts are at llmpricewatch.com. The tracker checks OpenRouter daily, so any new price moves on these models will show up there first.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.