China LLM API Token Plans 2026: Every Model, Every Plan, One Table

PJ • 2026-09-04 • China LLM Token Plan DeepSeek Qwen Kimi GLM Doubao ERNIE MiniMax Coding Plan API Pricing

LLM API prices in China change often. DeepSeek now uses peak and off-peak pricing. Kimi retired its K2.5 models on 2026-08-31. Old blog posts carry wrong numbers.

This page lists the major LLM APIs available in China. Every price comes from an official page. Every table row links to its source. When a number disagrees with your bill, follow the link and read the current page.

Last verified: 2026-09-09. The Update Log lists what changed and why.

Quick Pick: Which Plan Fits You

Your use case Pick Why
Coding agent, daily use GLM Coding Plan Lite ¥118/mo Credit caps (5-hour + weekly); 20+ tools
Coding agent, heavy volume GLM Coding Plan Pro ¥538/mo 6× Lite credits; same caps
Coding, request-based Volcano 方舟 Coding Plan 18,000 requests/mo (Lite); ≈10% of API cost; promo entry ¥9.9
Coding, cheapest entry Kimi K2.7 Code ¥49/mo Lowest entry price on this page; weekly quota
Heavy reasoning, cheapest per token DeepSeek V4-Flash, off-peak ¥1.5 in / ¥4.5 out per M; half the peak price
Long-context documents Qwen 3.8 Flash ¥0.8 in / ¥2.7 out; 1M context
Long-horizon flagship work Kimi K3 ¥20 in / ¥100 out; 1M context
Low-latency chat product Doubao Seed-2.1-turbo ¥3 in / ¥15 out; 500K free tokens
Batch jobs ERNIE 4.5 Turbo batch ¥0.32 in / ¥1.28 out; 60% off real-time

Run one coding subscription and one pay-as-you-go key. When the subscription quota runs out, the pay-as-you-go key takes over.

Pay-as-You-Go: Flagship Models

Prices per million tokens, RMB. MiniMax quotes USD. "In" = input, "Out" = output. Cache-hit input costs less where listed.

Model Platform In (miss) In (hit) Out Context Notes
DeepSeek V4-Flash (0731) api.deepseek.com ¥1.50 off-peak / ¥3.00 peak ¥0.05 / ¥0.10 ¥4.50 / ¥9.00 1M Peak = Mon-Fri 9:00-12:00, 14:00-18:00 Beijing
DeepSeek V4-Pro (0813) api.deepseek.com ¥4.50 / ¥9.00 ¥0.15 / ¥0.30 ¥13.50 / ¥27.00 1M Same peak window
Kimi K3 platform.kimi.com ¥20.00 ¥2.00 ¥100.00 1M Reasoning effort: low / high / max
Qwen 3.8 Max Aliyun Bailian ¥12.00 10% of input price ¥36.00 1M Batch = 50% of price
Qwen 3.8 Max Prime Aliyun Bailian ¥24.00 10% of input price ¥72.00 1M Fast lane, same context
GLM-5.3 SiliconFlow (hosts Zhipu) ¥8.00 ¥2.00 ¥28.00 Zhipu price page is JS-rendered; this is the SiliconFlow hosted rate
ERNIE 5.1 Baidu Qianfan ¥4.00 (≤32K) / ¥6.00 (32-128K) ¥18.00 / ¥22.00 128K No batch lane
ERNIE 4.5 Turbo Baidu Qianfan ¥0.80 ¥0.20 ¥3.20 128K Batch = 40% of price
Doubao Seed-2.1-pro Volcano Engine ¥6.00 ¥1.20 ¥30.00 Cache storage: ¥0.017/M/hour
Doubao Seed-2.1-turbo Volcano Engine ¥3.00 ¥0.60 ¥15.00 Cache storage: ¥0.017/M/hour
MiniMax M3 platform.minimax.io $0.30 (≤512K) / $0.60 (>512K) $0.06 $1.20 / $2.40 Permanent 50% off
MiniMax M2 (legacy) platform.minimax.io $0.30 $0.03 $1.20 Cache write $0.375/M

DeepSeek peak pricing, worked example

Off-peak: 1M input + 1M output = ¥1.50 + ¥4.50 = ¥6.00. Peak: ¥3.00 + ¥9.00 = ¥12.00. Doubled. Cache-hit input is 98.3% off peak-miss (¥0.05 vs ¥3.00). Put the system prompt first so it hits the cache. Run batch jobs after 18:00 Beijing.

GLM Coding Plan, subscription math

GLM publishes credit quotas, not flat token counts:

Plan 5-hour credits Weekly credits
Lite 2,000 10,000
Pro 12,000 60,000
Max 28,000 140,000

GLM-5.3 coefficients per 10K tokens: input 6.9, cache-hit 1.7, output 24. Formula: credits = (input × 6.9 + cache-hit × 1.7 + output × 24) / 10,000. Outside peak hours (Mon-Fri 14:00-18:00 UTC+8), calls cost 50% of the listed credits. GLM claims up to 92% savings versus the pay-as-you-go API at high cache-hit rates. Source.

GLM lists ¥118 / ¥538 / ¥1,078 per month for Lite / Pro / Max. Quarterly billing pays 80% of the monthly price (¥94.4 / ¥430.4 / ¥862.4). Source.

Coding Subscriptions (Monthly, Fixed)

Plan Platform Price What you get
GLM Coding Plan Lite / Pro / Max bigmodel.cn/glm-coding ¥118 / ¥538 / ¥1,078 monthly; 80% price quarterly 5-hour + weekly credit caps; GLM-5.3 + 5.3-Flash; 20+ tools (Claude Code, Cursor, Cline, Trae, OpenCode, Kilo)
Kimi K2.7 Code kimi.com ¥49 / ¥99 / ¥199 / ¥699 monthly Weekly quota; agent coding; includes K2.6 membership benefits
方舟 Coding Plan Lite / Pro Volcano Engine Promo entry from ¥9.9 (limited time) 18,000 requests/mo (Lite); ≈10% of raw API cost per official article
MiniMax Token Plan Plus / Max platform.minimaxi.com ¥49 / ¥119 monthly M2 access inside the MiniMax Code agent; individual tier

Workhorse Models (Cheap, Fast)

Model In Out Context Platform
Qwen Flash ¥0.15 ¥1.50 128K Bailian
Qwen 3.7 Flash ¥0.20 ¥0.80 32K Bailian
Qwen 3.8 Flash ¥0.80 ¥2.70 1M Bailian
Qwen Plus (2025-12) ¥0.80 ¥2.00 non-thinking / ¥8.00 thinking 128K Bailian
Qwen 3.7 Plus ¥2.00 (limited-time 80% = ¥1.60) ¥8.00 (= ¥6.40) 256K Bailian
Doubao Seed-2.1-turbo ¥3.00 ¥15.00 Volcano Engine
Kimi K2.7 Code (API) ¥6.50 ¥27.00 256K Kimi
Kimi K2.7 Code Highspeed ¥13.00 ¥54.00 256K Kimi
GLM-4.5-Air ¥1.00 ¥6.00 SiliconFlow
GLM-5.3-Flash ¥0.80 (derived, see note) ¥2.80 1M GLM docs
ERNIE 4.5 Turbo VL ¥3.00 ¥9.00 (batch ¥3.60) 32K Qianfan
LongCat 2.0 (Meituan) ¥5.00 ¥20.00 SiliconFlow
MiniMax M2.5 ¥2.10 ¥8.40 SiliconFlow

GLM-5.3-Flash note: Zhipu states that GLM-5.3-Flash costs 1/10 of GLM-5.3, and 1/20 during the limited-time discount. The ¥0.80 / ¥2.80 figures come from dividing the SiliconFlow GLM-5.3 rate (¥8 / ¥28) by 10. The quality comparison to GLM-5.3 is a vendor claim, not an independent benchmark. Source.

Free Grants (Register and Use)

Platform Grant Expiry / notes
Aliyun Bailian 1M tokens per Qwen model 90 days; North China 2 (Beijing) region only. Source
Volcano 方舟 500K tokens per Doubao model Seed-2.1-pro, Seed-2.1-turbo, Seed-Evolving. Source
Zhipu bigmodel.cn 20M tokens on signup Referral grants go up to 200M. Source

The Zhipu grant alone covers 20M tokens. Use the grants for experiments before you spend money.

The Three Pricing Rules That Matter

  1. Peak/off-peak splits the market. DeepSeek charges 2× at peak (Mon-Fri 9:00-12:00, 14:00-18:00 Beijing). GLM cuts credits to 50% outside Mon-Fri 14:00-18:00 UTC+8. Run batch jobs at night. The bill drops to half.
  2. Cache hits are the largest discount. Vendors price cache hits between 3% and 25% of miss prices: DeepSeek 3.3%, Kimi 10%, Qwen 10%, ERNIE 25%. Keep system prompts static and put them first.
  3. Batch lanes cost less. Baidu batch runs at 40% of the real-time price (60% off). Aliyun Batch runs at 50% of price. Use batch when latency is not a constraint.

Which One Do I Pick (Decision Flow)

How to Update This Article

Prices on this page move without notice. Run this checklist monthly, or after any vendor price announcement:

  1. DeepSeek — check peak windows and both model rows.
  2. Aliyun Bailian — Qwen Max / Plus / Flash tiers and the free-quota column.
  3. Kimi K3 pricing + K2.7 Code — monthly plan prices.
  4. GLM coding plan — credit coefficients change. bigmodel.cn/glm-coding for plan prices.
  5. Doubao product page — Seed model tiers and free grants.
  6. ERNIE pricing — batch discounts.
  7. MiniMax pay-as-you-go + Token Plan.
  8. SiliconFlow pricing — third-party hosting of Zhipu, DeepSeek, Kimi, and MiniMax models. Use it to catch price changes before the vendors publish them.

If a number fails re-verification against an official link, drop it from this page. This page keeps only sourced facts.

Update Log

Share this article
AI Tools Insight may earn a commission from some links. Editorial independence is always maintained.