cocodot
← Back to guides
Local card declined for overseas AI? cocodot: one card + one key
ExplainersUpdated 2026-06

How to Evaluate an AI API Gateway Before You Trust It With Money (2026)

Gateways that resell model access vary enormously in integrity. Five checks that separate a serious operator from one you will regret prepaying.

TL;DR: Gateways that resell access to Claude, GPT and Gemini differ far more in integrity than in price, and the differences are invisible from the marketing page. Five checks matter, roughly in this order. (1) Model integrity — does it actually serve the model it charges you for, or a cheaper substitute? This is the single most common failure and you cannot detect it by eye. (2) Counterparty risk — most gateways take prepaid balances, so a disclosed operating entity and published terms covering balances and suspension are not optional. (3) Model coverage — availability differs; the model you depend on may not be there. (4) Rate limits and concurrency — a cheap gateway that throttles you at production load is not cheap. (5) Support reachability — when something breaks at 2am, is there anyone. Price should be the last tiebreaker, not the first filter: the cheapest listing in this market is frequently the one substituting models.

1. Why price is the wrong first filter

Reselling model access has thin, fairly predictable margins: the operator buys capacity and marks it up. When a listing is dramatically below the rest of the market, that gap has to come from somewhere, and there are only a few places it can come from — substituting a cheaper model behind the scenes, running on capacity of uncertain provenance, or operating at a loss to accumulate prepaid balances. None of those is something you want to discover after wiring money. Price still matters, but as a tiebreaker among operators that passed the other checks.

2. Model integrity: the failure you cannot see

Model substitution is the defining risk of this market. A gateway bills you for a frontier model and quietly serves a cheaper one; outputs still look plausible, so the substitution goes unnoticed until quality problems surface in production. Eyeballing responses does not catch it — you need a repeatable test. The practical approach is a fixed prompt set that probes capabilities where models differ measurably, run against the endpoint, with results scored consistently. An open-source implementation exists (probe.cocodot.co): it sends a set of probes to any OpenAI-compatible endpoint and reports whether the responses are consistent with the claimed model. It works against any provider — including the one that publishes it, which is the point.

3. Counterparty risk is structural, not paranoid

Nearly every gateway operates on prepaid balances, which means at any moment you are an unsecured creditor of that operator. This is not a reason to avoid gateways; it is a reason to check who you are extending credit to. Look for: a named operating entity (not just a support handle), published terms that specifically address what happens to balances on account suspension or service termination, and a funding pattern that matches the stated model. A practical mitigation regardless of operator: keep prepaid balances at roughly the size of your near-term usage rather than buying large amounts for a discount.

4. Coverage and limits: check, do not assume

'All the major models' is a marketing claim, not a specification. Query the models endpoint yourself and confirm the exact model identifiers you depend on are present. Then check limits: request rate, tokens per minute, and concurrency. These are the parameters that decide whether a gateway survives your production traffic, and many operators do not publish them. If limits are undisclosed, test at your realistic concurrency before you commit — a gateway that throttles under load costs far more than the price difference that attracted you.

5. Support, and what a good answer sounds like

Before paying anything, send a specific technical question — about rate limits, or how they handle a particular error code. What comes back tells you a great deal. A serious operator answers concretely, acknowledges limitations, and can explain their upstream arrangement in general terms. A risky one answers with marketing copy, avoids specifics, or does not answer at all. You are not testing politeness; you are testing whether there is technical staff on the other end when something breaks.

6. Disclosure

cocodot, which publishes this page, operates one such gateway — so weigh this accordingly. Its specifics: OpenAI-compatible endpoint, per-token billing against a prepaid balance, sourced through a licensed cloud provider's official resale channel rather than opaque capacity, and model integrity independently checkable with the open-source probe linked above. It also issues US-BIN virtual cards for people who would rather pay the official consoles directly. The reason for publishing the verification tool openly is that a claim you can test is worth more than a claim you are asked to believe — including ours.

7. A reasonable evaluation sequence

Put together, an efficient order: shortlist two or three operators on model coverage; send each a specific technical question and see what comes back; fund the smallest amount each allows; run the integrity probe and a short load test at your real concurrency; only then compare price per token on the models you actually call. This takes an afternoon and costs very little, and it front-loads the discovery of problems that would otherwise surface in production.

Five checks, and how to actually perform each

CheckHow to verifyRed flag
Model integrityRun a fixed prompt set, score the responsesRefuses to discuss verification
Counterparty riskFind the operating entity and termsOnly a chat handle, no entity
Model coverageQuery the models endpoint yourselfClaims 'all models' without a list
Rate limitsLoad-test at your real concurrencyNo published limits
SupportSend a question before payingNo reply, or bot-only
PriceCompare per-token, same modelFar below everyone — ask why

FAQ

What is the biggest risk with an AI API gateway?

Model substitution — being billed for a frontier model while a cheaper one answers. It is invisible without a repeatable test, which is why running an integrity probe before committing matters more than comparing prices.

Is the cheapest gateway a bad sign?

Not automatically, but a price far below the rest of the market has to be explained by something. Ask where the margin comes from, and verify model integrity before assuming it is simply efficiency.

How much should I prepay?

Roughly your near-term usage. Prepaid balances make you an unsecured creditor of the operator; large discounted top-ups increase that exposure for a small saving.

How do I check model integrity myself?

Run a fixed prompt set against the endpoint and score consistency with the claimed model. An open-source probe (probe.cocodot.co) does this for any OpenAI-compatible endpoint, including cocodot's own.

Should I use a gateway or the official API?

Official if you need invoicing, contractual SLA, or a direct vendor relationship. A gateway if card payment is the blocker or you want several vendors behind one key — accepting that you are trusting the operator.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, $1 per active card per month; spending: $0.60 settlement fee on purchases under $20; a corresponding fee applies when the issuer charges one.

Service scope, pricing and limits →
How to Evaluate an AI API Gateway Before You Prepay