cocodot
← Back to guides
Local card declined for overseas AI? cocodot: one card + one key
AI APIUpdated 2026-09

Is Your API Relay Serving the Real Model? Substitution, Quantization and How to Verify (2026)

Paying flagship rates and quietly getting something lighter is two distinct problems — a swapped model and a quantized one fail differently. Here are the checks that separate them, and what the response metadata can and cannot prove.

TL;DR: "Downgrading" covers two different failures that need different tests. Substitution is a relay advertising a flagship and routing your request to a cheaper model — you pay flagship rates and, on easy prompts, never notice. Quantization is the same model served at reduced numerical precision: nothing was swapped, but output quality degrades subtly, usually late in long generations rather than on short answers. Substitution shows up as a cliff on hard problems; quantization shows up as drift. Four checks settle both in about ten minutes. ① Identity and context — ask what it is (weak evidence), then push a genuinely long document and see whether the advertised window is real. ② Capability — three to five genuinely hard prompts run against both a known-good endpoint and the candidate; the gap opens on hard problems, never on easy ones. ③ Latency, success rate and billing — 20–50 requests on one key, recording time-to-first-token, failures, truncations, and whether you were charged for the failures. ④ Response metadata — token counts are a better clue than the echoed model name, for reasons in section 6. If you would rather not build the harness, cocodot's checker is open source and hosted at probe.cocodot.co; paste any base URL and key. It works against us too, and a provider's willingness to be tested is worth more than any no-downgrade promise.

1. What downgrading actually means

A relay lists a flagship, charges accordingly, and routes some or all of that traffic to something cheaper. The economics are plain: flagship inference is genuinely expensive, and the margin between what you pay and what a small model costs is large enough to be tempting. Precision matters here, though, because not every substitution is dishonest. Reputable aggregators document fallback routing openly — when an upstream is at capacity they say so, and publish which provider served each request. That is a disclosed engineering trade-off, not a deception. The problem is the operators who do it silently, and the reason it persists is that on the prompts most people test with, the difference is invisible.

2. Substitution and quantization are different failures

These get conflated constantly, and conflating them means running the wrong test. Substitution means a different model entirely answered your request. Quantization means the advertised model ran, but with its weights stored at reduced numerical precision — int8 or int4 instead of the original format — which shrinks memory and cost at some accuracy cost. Quantization is a completely legitimate, widely used technique; serving a quantized model while billing for the full-precision one is what is not. The symptoms differ in a useful way. A substituted model falls off a cliff on genuinely hard problems while staying fine on easy ones. A quantized model drifts: short answers look right, and degradation concentrates in the tail of long generations — a long proof or a long code file where later sections quietly lose coherence with earlier ones, rare tokens get mangled, and numeric detail goes soft. So if your outputs are subtly worse but nothing obviously breaks, test with long generations, not with more hard questions. One honest caveat: from outside a black box you can rarely prove quantization. What you can do is compare against a known-good endpoint and see whether a gap reproduces.

3. Check one: identity and the context window

Ask the model to state what it is. Then treat the answer as weak evidence only — a system prompt can make any model introduce itself as a flagship, and this is the check people over-rely on. The harder test in the same family is the context window: send a document of tens of thousands of tokens and confirm it is ingested whole, without truncation and without an error. A cheaper substitute frequently cannot hold the window the listing advertises, and long input is exactly where that becomes visible. Make the document one you can ask a specific question about — something stated only in its final pages — so you are testing retrieval across the full window rather than just the absence of an error.

4. Check two: a benchmark that actually discriminates

Keep a small fixed set of three to five prompts with real discrimination between tiers: one long multi-step reasoning problem, one long code generation-and-debug task, one multi-step maths problem where a single arithmetic slip changes the answer. Run each against a known-good endpoint and against the candidate, and compare correctness and consistency. The discipline that makes this work: easy prompts prove nothing, because every model in the market answers them correctly, and a test everyone passes carries no information. Re-use the same set every time so results are comparable across providers and across months — a fixed benchmark you keep is worth far more than a better one you run once.

5. Check three: latency, failures, and what you were billed

Send 20–50 requests on one key and record time-to-first-token, the overall success rate, and any unexplained failures or truncations. This matters most in production, where intermittent timeouts and tokens charged on failed requests drain a budget without ever surfacing as an obvious incident. Two things to watch that people miss. First, unusually good latency is a signal, not a win — a smaller model is genuinely faster, so a flagship that responds suspiciously quickly deserves a capability check rather than congratulation. Second, reconcile the billing: add up the tokens in your own request log and compare against the provider's usage page for the same window. A persistent gap means either failed requests are being charged or the accounting is wrong, and both are worth knowing before you scale.

6. What the response metadata can and cannot tell you

Every OpenAI-compatible response carries structured fields, and they vary sharply in evidential value. The echoed model name is nearly worthless — the relay writes that string and can write anything. The token counts in the usage block are much more interesting, because different model families use different tokenizers: the same input text yields measurably different prompt_tokens under Anthropic's tokenizer than under OpenAI's. So send a fixed passage, record prompt_tokens, and compare against what the vendor's own tokenizer or API reports for that identical text. A consistent discrepancy points at a different model family in the path — and unlike the model name, this is not something a relay can trivially forge without also faking its billing. Two smaller tells: finish_reason and stop_reason enumerations differ between vendors, so values that do not belong to the advertised vendor's schema suggest a translation layer; and request-id formats are vendor-specific, so an id that does not match the upstream's shape indicates the response was re-serialised rather than passed through. None of these is conclusive alone. Together with the capability benchmark, they turn a suspicion into a well-founded one.

7. Use the open-source probe rather than building your own

cocodot published its downgrade checker as open source, with a hosted version at probe.cocodot.co. Paste in any OpenAI-compatible base URL and an API key and it runs the probes and reports what it found; the key is used for that request only and is not stored. The methodology is public specifically so that you can point it at us as readily as at anyone else, and so that you can read what it tests rather than taking a score on faith. Run it against your current provider before you renew, and against any candidate before you migrate — it takes a few minutes and it replaces an argument with a measurement.

8. Pick a provider that lets you test it

The decision rule that survives contact with reality is not "who promises not to downgrade" — every operator promises that, and the promise costs nothing to make. It is who publishes enough for you to check: model identity, real context window, per-model pricing, and an endpoint you can benchmark before committing. Fund a small amount, run the four checks, and scale only after your own numbers come back clean. If a probe does come back bad, the sequence is unglamorous but effective: re-run it to rule out a transient upstream fallback, ask the operator directly and read the answer for whether it engages with the specifics, stop adding balance, and migrate if the answer is evasive. That advice holds whichever provider you end up with, including this one.

Four checks: what to run, and what each failure looks like

CheckHow to run itA real flagshipSubstitutedQuantized
1. Identity + contextAsk what model it is; then send a very long documentConsistent identity, ingests the full windowVague identity, truncates or errors on long inputIdentity and window both fine
2. Capability benchmark3–5 hard prompts, same prompts against a known-good endpointSolves hard problems, stable styleFine on easy prompts, collapses on hard onesMostly right, degrades late in long outputs
3. Latency, failures, billing20–50 requests on one key; record TTFT, failures, chargesSteady TTFT and success rateOften unusually fast — smaller modelNormal-looking latency
4. Response metadataCompare usage token counts against the vendor's own tokenizerCounts line upCounts drift — different tokenizerCounts line up
Shortcutprobe.cocodot.co with any base_url + key (key not stored)Runs the probes and reports——

FAQ

How do I know if an API relay is serving the real model?

Run four checks: verify identity and push a very long context to see whether it truncates; run 3–5 hard prompts side by side against a known-good endpoint; send 20–50 requests and record latency, failures and what you were charged; and compare the usage token counts against the vendor's own tokenizer. If all four line up you are very likely getting what you paid for. Or run probe.cocodot.co, which does it for you.

What is a quantized LLM, and would I notice?

A model whose weights are stored at reduced numerical precision — int8 or int4 rather than the original format — which cuts memory and cost at some accuracy cost. The technique is legitimate and widespread; billing for full precision while serving quantized is not. You would rarely notice on short answers. Degradation concentrates in the tail of long generations, so test with long outputs rather than more hard questions.

Can I trust a model that tells me which model it is?

Only as a hint. A system prompt can make a small model introduce itself as a flagship, so identity always needs corroborating with the context-window and capability checks. A model can claim to be a flagship; it cannot fake solving flagship-level problems.

Does the response metadata prove which model ran?

The echoed model name proves nothing — the relay writes that string. The usage token counts are better evidence, because model families tokenize differently, so the same text produces different prompt_tokens under different tokenizers. Vendor-specific finish_reason values and request-id formats are smaller tells. None is conclusive alone; combined with a capability benchmark they are persuasive.

Does a very cheap relay always mean something is wrong?

Not necessarily, but it earns an extra round of testing. Flagship inference has a hard upstream cost, so pricing well below it generally comes from one of three places: silent substitution, resold credits of unclear origin, or collecting prepaid balances at a loss. Each of those eventually becomes your problem. Test with a small amount before concentrating spend anywhere.

The probe came back bad — what should I do?

In order: re-run it, since a one-off can be a disclosed capacity fallback rather than a policy; ask the operator directly and judge whether the reply engages with specifics or deflects; stop topping up while it is unresolved; and migrate if the answer is evasive. Keep your prepaid balance small enough that this decision is never expensive to make.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, $1 per active card per month; spending: $0.60 settlement fee on purchases under $20; a corresponding fee applies when the issuer charges one.

Service scope, pricing and limits →
Is Your API Relay Serving the Real Model? How to Verify