cocodot
← Back to guides
Local card declined for overseas AI? cocodot: one card + one key
ChoosingUpdated 2026-08

Official API or Relay? Four Scenarios — Official Wins One of Them (2026)

Relays are not always cheaper. The item most often miscalculated is prompt caching: the catch is not whether it hits but how each layer prices cached input, and on that one line going direct to the official API is cheaper. Output-heavy and long-context workloads usually favor relays.

TL;DR: Do not believe 'relays are always cheaper'. Four scenarios: ① high cache hit rate (agents that resend the same long system prompt every turn) — official is usually cheaper, and not because relays fail to hit: caching works here. The difference is pricing. Official bills cached input at 1/10 to 1/30 of normal input; we bill it at half our input rate, so the higher your hit rate and the more input-heavy you are, the more official wins. ② Output-dominated — relays are usually cheaper; there is no caching on the output side. ③ Long context — check whether the relay publishes tiered pricing; be wary if not. ④ Official price increases or time-of-day pricing — a relay's cost comes from its own upstream contracts and does not necessarily move in step, so the gap can widen suddenly (and may close later, so compute per period). To know which you are, look at one thing: how much of each request is a prefix that never changes.

1. Admit it first: relays are not always cheaper

Articles in this industry almost uniformly claim relays beat official on price, but that holds only in some scenarios. In others, official is clearly better value — not because a relay is badly run, but by structure. Saying so plainly is more useful than listing three more advantages, because you will work it out yourself eventually.

2. The key dividing line: prompt caching

Mainstream models now offer prompt caching: send the same prefix a second time and the cached portion is charged at a much lower rate — roughly 1/10 of normal input for Claude, and as low as about 1/30 for DeepSeek (V4 family: cache hit ¥0.05 versus miss ¥1.5). Who benefits most? Agent applications, which resend the system prompt, tool definitions and history every turn with a nearly identical prefix. Across 50 turns, the last 49 prefixes could cost 1/10.

3. Whether caching hits depends on routing — how to test it yourself

Caches are bound to a specific upstream account or instance: the first request's cache lands on account A, and if the second routes to account B the prefix misses. So whether a relay hits depends on how it routes, and that varies a lot between vendors — which is exactly the thing worth measuring yourself instead of taking anyone's word. The test is simple: send the same long prefix three times and read the cache fields in usage, checking whether the second and third jump (a copy-paste script is in our piece on measuring prompt caching). On cocodot specifically, both cache reads and writes work, cached input is billed at half the input rate, and usage returns the cache token counts so you can verify it — hit rate depends on how stable your prefix is, and we do not promise a number.

4. No such problem on the output side

Caching applies only to input. Output tokens are full price on every route. If your usage is output-heavy — long-form generation, translation, batch rewriting, content production — caching is not a variable, and whoever has the lower output unit price wins; relays usually have the edge there. Quick self-check: open your usage stats and look at the input-to-output ratio. Input a dozen times output → you are agent-shaped and should take caching seriously; output-heavy → compare unit prices directly.

5. Long context: check for published tiered pricing

Several models price input by length tier, rising with length (GPT-5.5, for example, doubles the unit price past 272K input). Vendors publish the tiers. If a relay quotes only the lowest tier, you budget low and get charged high, and the bill makes no sense. Ask outright: where can I see the tiers? If you cannot, assume a trap. cocodot lists every tier on its pricing page (cocodot.co/pricing#models, no sign-up needed) — no hidden long-context markup.

6. When official reprices, the gap can widen suddenly

DeepSeek enabled peak/off-peak pricing from 2026-08-17 00:00 Beijing time: peak hours are 9:00–12:00 and 14:00–18:00 Beijing time, everything else is off-peak at half the peak price. For V4-Flash, input (cache miss) is ¥1.5 off-peak / ¥3.0 peak and output ¥4.5 / ¥9.0 — all clearly above the prior level. Such changes hit relays with a lag: a relay's cost comes from its own upstream contracts and does not change because the vendor's retail price did — though it may follow later. So every official repricing deserves a fresh calculation; the earlier conclusion may have flipped. That is why 'which is cheaper' has no permanent answer, only a current one.

7. So who is cocodot for

A fit if: you mainly generate output, or you cannot pay official from China at all (Claude, GPT and Gemini accept overseas cards only), or you want one key and one balance across several vendors. Cache-heavy agent loops are cheaper on the official API, as the table above says — unless you cannot pay official from China at all. Per-model prices are public at cocodot.co/pricing#models without sign-up, so run the numbers yourself. Trying it is cheap: sign up and verify email for a $0.5 trial credit, set base_url to https://cocodot.co/api/ai/v1 and your key, use official model names or codes, change nothing else, and a few calls will show whether latency and quality hold up.

Which wins in four scenarios (2026-08)

Your scenarioTypical exampleCheaper option
High cache hit rateAgent loops resending large system prompts and tool definitions every turnOfficial — it bills cached input at 1/10 of its list price; we bill half of our own input rate
Output-dominatedLong-form generation, translation, batch rewritingRelay
Long contextWhole codebase / long documents in contextDepends on whether the relay publishes tiered pricing
Official price increase / time-of-day pricingDeepSeek peak/off-peak pricing from 2026-08-17 00:00 Beijing timeRelay — cost structure does not necessarily follow

FAQ

How do I know whether my cache hit rate is high?

Look at request structure: a large fixed system prompt, tool definitions or knowledge prefix on every call (typical agents, support bots, code assistants) means a high hit rate. If every input differs (translation, classification, single-turn Q&A), there is essentially nothing to cache.

Can the relay caching problem be fixed?

Yes, but the relay must implement session stickiness — routing one API key's requests to the same upstream backend. It is engineering work and not every relay has done it. Asking before you sign up is the easy way to find out.

Why write about scenarios where you lose?

Because you would work it out eventually. Better we say it than you discover it on a bill — and it makes our 'we are cheaper here' claims a little more credible.

Will your prices rise when official reprices?

Not necessarily in step. Relay cost comes from upstream contracts, a separate line from official retail. After any official repricing, recheck — our live per-model prices are at cocodot.co/pricing#models, or fetch GET https://cocodot.co/api/ai/models directly.

Can I use official and a relay at the same time?

Yes, and many do: cache-heavy agents on official, batch output and models you cannot pay for from China on the relay. Both are OpenAI-compatible; switch base_url per scenario.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, $1 per active card per month; spending: $0.60 settlement fee on purchases under $20; a corresponding fee applies when the issuer charges one.

Service scope, pricing and limits →
Official API or Relay? Four Scenarios, Official Wins One