LLM API prices in China change often. DeepSeek now uses peak and off-peak pricing. Kimi retired its K2.5 models on 2026-08-31. Old blog posts carry wrong numbers.
This page lists the major LLM APIs available in China. Every price comes from an official page. Every table row links to its source. When a number disagrees with your bill, follow the link and read the current page.
Last verified: 2026-09-09. The Update Log lists what changed and why.
Quick Pick: Which Plan Fits You
| Your use case | Pick | Why |
|---|---|---|
| Coding agent, daily use | GLM Coding Plan Lite ¥118/mo | Credit caps (5-hour + weekly); 20+ tools |
| Coding agent, heavy volume | GLM Coding Plan Pro ¥538/mo | 6× Lite credits; same caps |
| Coding, request-based | Volcano 方舟 Coding Plan | 18,000 requests/mo (Lite); ≈10% of API cost; promo entry ¥9.9 |
| Coding, cheapest entry | Kimi K2.7 Code ¥49/mo | Lowest entry price on this page; weekly quota |
| Heavy reasoning, cheapest per token | DeepSeek V4-Flash, off-peak | ¥1.5 in / ¥4.5 out per M; half the peak price |
| Long-context documents | Qwen 3.8 Flash | ¥0.8 in / ¥2.7 out; 1M context |
| Long-horizon flagship work | Kimi K3 | ¥20 in / ¥100 out; 1M context |
| Low-latency chat product | Doubao Seed-2.1-turbo | ¥3 in / ¥15 out; 500K free tokens |
| Batch jobs | ERNIE 4.5 Turbo batch | ¥0.32 in / ¥1.28 out; 60% off real-time |
Run one coding subscription and one pay-as-you-go key. When the subscription quota runs out, the pay-as-you-go key takes over.
Pay-as-You-Go: Flagship Models
Prices per million tokens, RMB. MiniMax quotes USD. "In" = input, "Out" = output. Cache-hit input costs less where listed.
| Model | Platform | In (miss) | In (hit) | Out | Context | Notes |
|---|---|---|---|---|---|---|
| DeepSeek V4-Flash (0731) | api.deepseek.com | ¥1.50 off-peak / ¥3.00 peak | ¥0.05 / ¥0.10 | ¥4.50 / ¥9.00 | 1M | Peak = Mon-Fri 9:00-12:00, 14:00-18:00 Beijing |
| DeepSeek V4-Pro (0813) | api.deepseek.com | ¥4.50 / ¥9.00 | ¥0.15 / ¥0.30 | ¥13.50 / ¥27.00 | 1M | Same peak window |
| Kimi K3 | platform.kimi.com | ¥20.00 | ¥2.00 | ¥100.00 | 1M | Reasoning effort: low / high / max |
| Qwen 3.8 Max | Aliyun Bailian | ¥12.00 | 10% of input price | ¥36.00 | 1M | Batch = 50% of price |
| Qwen 3.8 Max Prime | Aliyun Bailian | ¥24.00 | 10% of input price | ¥72.00 | 1M | Fast lane, same context |
| GLM-5.3 | SiliconFlow (hosts Zhipu) | ¥8.00 | ¥2.00 | ¥28.00 | — | Zhipu price page is JS-rendered; this is the SiliconFlow hosted rate |
| ERNIE 5.1 | Baidu Qianfan | ¥4.00 (≤32K) / ¥6.00 (32-128K) | — | ¥18.00 / ¥22.00 | 128K | No batch lane |
| ERNIE 4.5 Turbo | Baidu Qianfan | ¥0.80 | ¥0.20 | ¥3.20 | 128K | Batch = 40% of price |
| Doubao Seed-2.1-pro | Volcano Engine | ¥6.00 | ¥1.20 | ¥30.00 | — | Cache storage: ¥0.017/M/hour |
| Doubao Seed-2.1-turbo | Volcano Engine | ¥3.00 | ¥0.60 | ¥15.00 | — | Cache storage: ¥0.017/M/hour |
| MiniMax M3 | platform.minimax.io | $0.30 (≤512K) / $0.60 (>512K) | $0.06 | $1.20 / $2.40 | — | Permanent 50% off |
| MiniMax M2 (legacy) | platform.minimax.io | $0.30 | $0.03 | $1.20 | — | Cache write $0.375/M |
DeepSeek peak pricing, worked example
Off-peak: 1M input + 1M output = ¥1.50 + ¥4.50 = ¥6.00. Peak: ¥3.00 + ¥9.00 = ¥12.00. Doubled. Cache-hit input is 98.3% off peak-miss (¥0.05 vs ¥3.00). Put the system prompt first so it hits the cache. Run batch jobs after 18:00 Beijing.
GLM Coding Plan, subscription math
GLM publishes credit quotas, not flat token counts:
| Plan | 5-hour credits | Weekly credits |
|---|---|---|
| Lite | 2,000 | 10,000 |
| Pro | 12,000 | 60,000 |
| Max | 28,000 | 140,000 |
GLM-5.3 coefficients per 10K tokens: input 6.9, cache-hit 1.7, output 24. Formula: credits = (input × 6.9 + cache-hit × 1.7 + output × 24) / 10,000. Outside peak hours (Mon-Fri 14:00-18:00 UTC+8), calls cost 50% of the listed credits. GLM claims up to 92% savings versus the pay-as-you-go API at high cache-hit rates. Source.
GLM lists ¥118 / ¥538 / ¥1,078 per month for Lite / Pro / Max. Quarterly billing pays 80% of the monthly price (¥94.4 / ¥430.4 / ¥862.4). Source.
Coding Subscriptions (Monthly, Fixed)
| Plan | Platform | Price | What you get |
|---|---|---|---|
| GLM Coding Plan Lite / Pro / Max | bigmodel.cn/glm-coding | ¥118 / ¥538 / ¥1,078 monthly; 80% price quarterly | 5-hour + weekly credit caps; GLM-5.3 + 5.3-Flash; 20+ tools (Claude Code, Cursor, Cline, Trae, OpenCode, Kilo) |
| Kimi K2.7 Code | kimi.com | ¥49 / ¥99 / ¥199 / ¥699 monthly | Weekly quota; agent coding; includes K2.6 membership benefits |
| 方舟 Coding Plan Lite / Pro | Volcano Engine | Promo entry from ¥9.9 (limited time) | 18,000 requests/mo (Lite); ≈10% of raw API cost per official article |
| MiniMax Token Plan Plus / Max | platform.minimaxi.com | ¥49 / ¥119 monthly | M2 access inside the MiniMax Code agent; individual tier |
Workhorse Models (Cheap, Fast)
| Model | In | Out | Context | Platform |
|---|---|---|---|---|
| Qwen Flash | ¥0.15 | ¥1.50 | 128K | Bailian |
| Qwen 3.7 Flash | ¥0.20 | ¥0.80 | 32K | Bailian |
| Qwen 3.8 Flash | ¥0.80 | ¥2.70 | 1M | Bailian |
| Qwen Plus (2025-12) | ¥0.80 | ¥2.00 non-thinking / ¥8.00 thinking | 128K | Bailian |
| Qwen 3.7 Plus | ¥2.00 (limited-time 80% = ¥1.60) | ¥8.00 (= ¥6.40) | 256K | Bailian |
| Doubao Seed-2.1-turbo | ¥3.00 | ¥15.00 | — | Volcano Engine |
| Kimi K2.7 Code (API) | ¥6.50 | ¥27.00 | 256K | Kimi |
| Kimi K2.7 Code Highspeed | ¥13.00 | ¥54.00 | 256K | Kimi |
| GLM-4.5-Air | ¥1.00 | ¥6.00 | — | SiliconFlow |
| GLM-5.3-Flash | ¥0.80 (derived, see note) | ¥2.80 | 1M | GLM docs |
| ERNIE 4.5 Turbo VL | ¥3.00 | ¥9.00 (batch ¥3.60) | 32K | Qianfan |
| LongCat 2.0 (Meituan) | ¥5.00 | ¥20.00 | — | SiliconFlow |
| MiniMax M2.5 | ¥2.10 | ¥8.40 | — | SiliconFlow |
GLM-5.3-Flash note: Zhipu states that GLM-5.3-Flash costs 1/10 of GLM-5.3, and 1/20 during the limited-time discount. The ¥0.80 / ¥2.80 figures come from dividing the SiliconFlow GLM-5.3 rate (¥8 / ¥28) by 10. The quality comparison to GLM-5.3 is a vendor claim, not an independent benchmark. Source.
Free Grants (Register and Use)
| Platform | Grant | Expiry / notes |
|---|---|---|
| Aliyun Bailian | 1M tokens per Qwen model | 90 days; North China 2 (Beijing) region only. Source |
| Volcano 方舟 | 500K tokens per Doubao model | Seed-2.1-pro, Seed-2.1-turbo, Seed-Evolving. Source |
| Zhipu bigmodel.cn | 20M tokens on signup | Referral grants go up to 200M. Source |
The Zhipu grant alone covers 20M tokens. Use the grants for experiments before you spend money.
The Three Pricing Rules That Matter
- Peak/off-peak splits the market. DeepSeek charges 2× at peak (Mon-Fri 9:00-12:00, 14:00-18:00 Beijing). GLM cuts credits to 50% outside Mon-Fri 14:00-18:00 UTC+8. Run batch jobs at night. The bill drops to half.
- Cache hits are the largest discount. Vendors price cache hits between 3% and 25% of miss prices: DeepSeek 3.3%, Kimi 10%, Qwen 10%, ERNIE 25%. Keep system prompts static and put them first.
- Batch lanes cost less. Baidu batch runs at 40% of the real-time price (60% off). Aliyun Batch runs at 50% of price. Use batch when latency is not a constraint.
Which One Do I Pick (Decision Flow)
- You run a coding agent daily → GLM Coding Plan Lite. Move to Pro when you use the 10,000 weekly credits.
- You want the cheapest general-purpose API → DeepSeek V4-Flash, off-peak, cache on.
- You process long documents daily → Qwen 3.8 Flash.
- You need hard reasoning in Chinese → Kimi K3. Accept the ¥100/M output price.
- You build a product with tight cost control → MiniMax M2 (¥8.4/M output) or Doubao Seed-2.1-turbo (¥15/M output).
- You only need one-off tasks → use the Aliyun or Volcano free grants (1M and 500K tokens per model).
How to Update This Article
Prices on this page move without notice. Run this checklist monthly, or after any vendor price announcement:
- DeepSeek — check peak windows and both model rows.
- Aliyun Bailian — Qwen Max / Plus / Flash tiers and the free-quota column.
- Kimi K3 pricing + K2.7 Code — monthly plan prices.
- GLM coding plan — credit coefficients change. bigmodel.cn/glm-coding for plan prices.
- Doubao product page — Seed model tiers and free grants.
- ERNIE pricing — batch discounts.
- MiniMax pay-as-you-go + Token Plan.
- SiliconFlow pricing — third-party hosting of Zhipu, DeepSeek, Kimi, and MiniMax models. Use it to catch price changes before the vendors publish them.
If a number fails re-verification against an official link, drop it from this page. This page keeps only sourced facts.
Update Log
- 2026-09-09 — Challenge pass. Changes:
- ERNIE 4.5 Turbo output corrected to ¥3.20/M (batch ¥1.28). VL output corrected to ¥9.00/M (batch ¥3.60). Source: Baidu Qianfan docs table.
- GLM plan prices moved from "reported" to official: ¥118 / ¥538 / ¥1,078 monthly, 80% quarterly. Source: bigmodel.cn/glm-coding.
- Kimi K2.7 Code context corrected to 256K (262,144 tokens). Source: official docs.
- Added MiniMax M3 row ($0.30 / $1.20, permanent 50% off). M2 marked legacy. Source: MiniMax official docs.
- Added ERNIE 5.1 row (¥4 / ¥18 ≤32K). Source: Baidu Qianfan docs.
- Removed unsourced rows: DeepSeek 5M grant, Baidu 1M grant, SiliconFlow 20M grant.
- Fixed batch wording: Baidu batch = 40% of price (60% off), not "40% off".
- Cache-hit range corrected to 3-25% of miss price (was "5-10%").
- DeepSeek cache example corrected to 98.3% off peak-miss.
- 2026-09-04 — Initial table. All figures pulled from the official pages listed above on this date. DeepSeek peak/off-peak split included. GLM plan prices were third-party reported and are now official (2026-09-09).