cocodot
← Back to guides
Local card declined, direct access hard? cocodot does both
Kimi K3Updated 2026-08

Kimi K3, GLM-5.3, Qwen3.8-Max: China's New Flagship LLM APIs Compared (2026)

Three new Chinese flagship models, all with 1M-token context. Real API prices, how they differ from the previous generation, and how to call them without a Chinese enterprise account.

TL;DR: China's flagship LLMs turned over a generation in Jul-Aug 2026, and all three new models ship **1M-token context**: **Kimi K3** (¥20/¥100 per M tokens, Moonshot's flagship for long-horizon coding and knowledge work), **GLM-5.3** (¥8/¥28, Zhipu's MoE flagship with context caching at ¥2/M on hits), **Qwen3.8-Max** (¥12/¥36, 2.4T parameters, positioned for autonomous multi-day coding projects). cocodot serves all three **at official list price with no markup**, while previous-generation models (K2.6 / GLM-5.2 / Qwen3.7-Max) carry 5-12% discounts. The difference is access: official platforms require Chinese enterprise verification and prepayment; here a few dollars via Alipay or card gets you testing, one key across all three vendors — plus Claude/GPT on the same key for side-by-side evals.

The generational shift: 1M context is now standard

All three new flagships ship 1M-token context — an entire repo, hundreds of pages, a full season of scripts in one call. Workflows that required chunking and RAG can now often just send the whole thing. The trade-off is top-tier pricing, so the first question is which tasks deserve a flagship, not which flagship is strongest.

What each one is for

**Kimi K3** targets long-horizon coding and end-to-end knowledge work — the most expensive of the three (¥100/M output), built for genuinely long agent chains. **GLM-5.3** is the value play at roughly a third of K3's price, with context caching (¥2/M on cache hits) that matters enormously for agents resending large system prompts. **Qwen3.8-Max** is a 2.4-trillion-parameter model positioned for autonomous multi-day project delivery and professional domains like law and finance.

Budget-conscious? The previous generation is closer than the price gap suggests

Honestly: for most everyday tasks the gap between generations is smaller than the price gap. K2.6's output price is a quarter of K3's. A sane default: run the previous generation with caching, escalate to the flagship only for the hard problems — same key, just change the model name.

Getting started without a Chinese enterprise account

Official platforms each require registration, verification and prepayment inside China. cocodot is OpenAI-compatible: point base_url to https://cocodot.co/api/ai/v1 and set model to kimi-k3 / glm-5.3 / qwen3.8-max (official names auto-route too). Email verification grants $0.5 trial credit for the budget tiers; top up to unlock everything. Live prices at /pricing — if this article ever disagrees with the pricing page, the pricing page wins.

New flagships vs previous generation: per million tokens, input/output

VendorNew flagship (list)Previous gen (discounted)Context
MoonshotKimi K3: ¥20/¥100K2.6: ¥6.17/¥25.65 (5% off)1M / 262K
ZhipuGLM-5.3: ¥8/¥28GLM-5.2: ¥7.6/¥26.6 (5% off)1M / 1M
AlibabaQwen3.8-Max: ¥12/¥36Qwen3.7-Max: ¥11.4/¥34.2 (5% off)1M / 1M

FAQ

Why are the new flagships at list price while older models get discounts?

Channel discounts for just-released models take time to negotiate. We price on our real cost — no markup, no fake discounts. When a discount lands, the pricing page updates immediately.

Are these the real models or degraded clones?

Official relay via licensed cloud vendors. Our verification tool is open source: probe.cocodot.co — test any relay, including us.

Can I mix these with Claude/GPT?

One key, one balance, 38 models across Claude/GPT/Gemini/DeepSeek and every major Chinese vendor — ideal for side-by-side evals.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, 0% on spend, $1 per active card per month.

Service scope, pricing and limits →
Kimi K3, GLM-5.3, Qwen3.8-Max: China's New Flagship LLM APIs Compared (2026) · cocodot