Is Your API Relay Serving the Real Model? Substitution, Quantization and How to Verify (2026)
Paying flagship rates and quietly getting something lighter is two distinct problems — a swapped model and a quantized one fail differently. Here are the checks that separate them, and what the response metadata can and cannot prove.
1. What downgrading actually means
A relay lists a flagship, charges accordingly, and routes some or all of that traffic to something cheaper. The economics are plain: flagship inference is genuinely expensive, and the margin between what you pay and what a small model costs is large enough to be tempting. Precision matters here, though, because not every substitution is dishonest. Reputable aggregators document fallback routing openly — when an upstream is at capacity they say so, and publish which provider served each request. That is a disclosed engineering trade-off, not a deception. The problem is the operators who do it silently, and the reason it persists is that on the prompts most people test with, the difference is invisible.
2. Substitution and quantization are different failures
These get conflated constantly, and conflating them means running the wrong test. Substitution means a different model entirely answered your request. Quantization means the advertised model ran, but with its weights stored at reduced numerical precision — int8 or int4 instead of the original format — which shrinks memory and cost at some accuracy cost. Quantization is a completely legitimate, widely used technique; serving a quantized model while billing for the full-precision one is what is not. The symptoms differ in a useful way. A substituted model falls off a cliff on genuinely hard problems while staying fine on easy ones. A quantized model drifts: short answers look right, and degradation concentrates in the tail of long generations — a long proof or a long code file where later sections quietly lose coherence with earlier ones, rare tokens get mangled, and numeric detail goes soft. So if your outputs are subtly worse but nothing obviously breaks, test with long generations, not with more hard questions. One honest caveat: from outside a black box you can rarely prove quantization. What you can do is compare against a known-good endpoint and see whether a gap reproduces.
3. Check one: identity and the context window
Ask the model to state what it is. Then treat the answer as weak evidence only — a system prompt can make any model introduce itself as a flagship, and this is the check people over-rely on. The harder test in the same family is the context window: send a document of tens of thousands of tokens and confirm it is ingested whole, without truncation and without an error. A cheaper substitute frequently cannot hold the window the listing advertises, and long input is exactly where that becomes visible. Make the document one you can ask a specific question about — something stated only in its final pages — so you are testing retrieval across the full window rather than just the absence of an error.
4. Check two: a benchmark that actually discriminates
Keep a small fixed set of three to five prompts with real discrimination between tiers: one long multi-step reasoning problem, one long code generation-and-debug task, one multi-step maths problem where a single arithmetic slip changes the answer. Run each against a known-good endpoint and against the candidate, and compare correctness and consistency. The discipline that makes this work: easy prompts prove nothing, because every model in the market answers them correctly, and a test everyone passes carries no information. Re-use the same set every time so results are comparable across providers and across months — a fixed benchmark you keep is worth far more than a better one you run once.
5. Check three: latency, failures, and what you were billed
Send 20–50 requests on one key and record time-to-first-token, the overall success rate, and any unexplained failures or truncations. This matters most in production, where intermittent timeouts and tokens charged on failed requests drain a budget without ever surfacing as an obvious incident. Two things to watch that people miss. First, unusually good latency is a signal, not a win — a smaller model is genuinely faster, so a flagship that responds suspiciously quickly deserves a capability check rather than congratulation. Second, reconcile the billing: add up the tokens in your own request log and compare against the provider's usage page for the same window. A persistent gap means either failed requests are being charged or the accounting is wrong, and both are worth knowing before you scale.
6. What the response metadata can and cannot tell you
Every OpenAI-compatible response carries structured fields, and they vary sharply in evidential value. The echoed model name is nearly worthless — the relay writes that string and can write anything. The token counts in the usage block are much more interesting, because different model families use different tokenizers: the same input text yields measurably different prompt_tokens under Anthropic's tokenizer than under OpenAI's. So send a fixed passage, record prompt_tokens, and compare against what the vendor's own tokenizer or API reports for that identical text. A consistent discrepancy points at a different model family in the path — and unlike the model name, this is not something a relay can trivially forge without also faking its billing. Two smaller tells: finish_reason and stop_reason enumerations differ between vendors, so values that do not belong to the advertised vendor's schema suggest a translation layer; and request-id formats are vendor-specific, so an id that does not match the upstream's shape indicates the response was re-serialised rather than passed through. None of these is conclusive alone. Together with the capability benchmark, they turn a suspicion into a well-founded one.
7. Use the open-source probe rather than building your own
cocodot published its downgrade checker as open source, with a hosted version at probe.cocodot.co. Paste in any OpenAI-compatible base URL and an API key and it runs the probes and reports what it found; the key is used for that request only and is not stored. The methodology is public specifically so that you can point it at us as readily as at anyone else, and so that you can read what it tests rather than taking a score on faith. Run it against your current provider before you renew, and against any candidate before you migrate — it takes a few minutes and it replaces an argument with a measurement.
8. Pick a provider that lets you test it
The decision rule that survives contact with reality is not "who promises not to downgrade" — every operator promises that, and the promise costs nothing to make. It is who publishes enough for you to check: model identity, real context window, per-model pricing, and an endpoint you can benchmark before committing. Fund a small amount, run the four checks, and scale only after your own numbers come back clean. If a probe does come back bad, the sequence is unglamorous but effective: re-run it to rule out a transient upstream fallback, ask the operator directly and read the answer for whether it engages with the specifics, stop adding balance, and migrate if the answer is evasive. That advice holds whichever provider you end up with, including this one.
Four checks: what to run, and what each failure looks like
| Check | How to run it | A real flagship | Substituted | Quantized |
|---|---|---|---|---|
| 1. Identity + context | Ask what model it is; then send a very long document | Consistent identity, ingests the full window | Vague identity, truncates or errors on long input | Identity and window both fine |
| 2. Capability benchmark | 3–5 hard prompts, same prompts against a known-good endpoint | Solves hard problems, stable style | Fine on easy prompts, collapses on hard ones | Mostly right, degrades late in long outputs |
| 3. Latency, failures, billing | 20–50 requests on one key; record TTFT, failures, charges | Steady TTFT and success rate | Often unusually fast — smaller model | Normal-looking latency |
| 4. Response metadata | Compare usage token counts against the vendor's own tokenizer | Counts line up | Counts drift — different tokenizer | Counts line up |
| Shortcut | probe.cocodot.co with any base_url + key (key not stored) | Runs the probes and reports | — | — |