Official API or Relay? Four Scenarios — Official Wins One of Them (2026)
Relays are not always cheaper. The item most often miscalculated is prompt caching: the catch is not whether it hits but how each layer prices cached input, and on that one line going direct to the official API is cheaper. Output-heavy and long-context workloads usually favor relays.
1. Admit it first: relays are not always cheaper
Articles in this industry almost uniformly claim relays beat official on price, but that holds only in some scenarios. In others, official is clearly better value — not because a relay is badly run, but by structure. Saying so plainly is more useful than listing three more advantages, because you will work it out yourself eventually.
2. The key dividing line: prompt caching
Mainstream models now offer prompt caching: send the same prefix a second time and the cached portion is charged at a much lower rate — roughly 1/10 of normal input for Claude, and as low as about 1/30 for DeepSeek (V4 family: cache hit ¥0.05 versus miss ¥1.5). Who benefits most? Agent applications, which resend the system prompt, tool definitions and history every turn with a nearly identical prefix. Across 50 turns, the last 49 prefixes could cost 1/10.
3. Whether caching hits depends on routing — how to test it yourself
Caches are bound to a specific upstream account or instance: the first request's cache lands on account A, and if the second routes to account B the prefix misses. So whether a relay hits depends on how it routes, and that varies a lot between vendors — which is exactly the thing worth measuring yourself instead of taking anyone's word. The test is simple: send the same long prefix three times and read the cache fields in usage, checking whether the second and third jump (a copy-paste script is in our piece on measuring prompt caching). On cocodot specifically, both cache reads and writes work, cached input is billed at half the input rate, and usage returns the cache token counts so you can verify it — hit rate depends on how stable your prefix is, and we do not promise a number.
4. No such problem on the output side
Caching applies only to input. Output tokens are full price on every route. If your usage is output-heavy — long-form generation, translation, batch rewriting, content production — caching is not a variable, and whoever has the lower output unit price wins; relays usually have the edge there. Quick self-check: open your usage stats and look at the input-to-output ratio. Input a dozen times output → you are agent-shaped and should take caching seriously; output-heavy → compare unit prices directly.
5. Long context: check for published tiered pricing
Several models price input by length tier, rising with length (GPT-5.5, for example, doubles the unit price past 272K input). Vendors publish the tiers. If a relay quotes only the lowest tier, you budget low and get charged high, and the bill makes no sense. Ask outright: where can I see the tiers? If you cannot, assume a trap. cocodot lists every tier on its pricing page (cocodot.co/pricing#models, no sign-up needed) — no hidden long-context markup.
6. When official reprices, the gap can widen suddenly
DeepSeek enabled peak/off-peak pricing from 2026-08-17 00:00 Beijing time: peak hours are 9:00–12:00 and 14:00–18:00 Beijing time, everything else is off-peak at half the peak price. For V4-Flash, input (cache miss) is ¥1.5 off-peak / ¥3.0 peak and output ¥4.5 / ¥9.0 — all clearly above the prior level. Such changes hit relays with a lag: a relay's cost comes from its own upstream contracts and does not change because the vendor's retail price did — though it may follow later. So every official repricing deserves a fresh calculation; the earlier conclusion may have flipped. That is why 'which is cheaper' has no permanent answer, only a current one.
7. So who is cocodot for
A fit if: you mainly generate output, or you cannot pay official from China at all (Claude, GPT and Gemini accept overseas cards only), or you want one key and one balance across several vendors. Cache-heavy agent loops are cheaper on the official API, as the table above says — unless you cannot pay official from China at all. Per-model prices are public at cocodot.co/pricing#models without sign-up, so run the numbers yourself. Trying it is cheap: sign up and verify email for a $0.5 trial credit, set base_url to https://cocodot.co/api/ai/v1 and your key, use official model names or codes, change nothing else, and a few calls will show whether latency and quality hold up.
Which wins in four scenarios (2026-08)
| Your scenario | Typical example | Cheaper option |
|---|---|---|
| High cache hit rate | Agent loops resending large system prompts and tool definitions every turn | Official — it bills cached input at 1/10 of its list price; we bill half of our own input rate |
| Output-dominated | Long-form generation, translation, batch rewriting | Relay |
| Long context | Whole codebase / long documents in context | Depends on whether the relay publishes tiered pricing |
| Official price increase / time-of-day pricing | DeepSeek peak/off-peak pricing from 2026-08-17 00:00 Beijing time | Relay — cost structure does not necessarily follow |