cocodot
← Back to guides
Local card declined for overseas AI? cocodot: one card + one key
AI APIUpdated 2026-08

Cutting AI API Costs: Model Tiering, max_tokens, and Eliminating Wasted Calls (2026 Guide)

API bill creeping up? Five fixes ranked by savings: model tiering (biggest), slimmer context, capped output, no wasted calls, watching the ledger.

TL;DR: When an API bill runs away, the reason is almost never 'too much usage' but expensive models doing cheap work — the number-one money pit; top-tier and Chinese models differ in unit price by more than an order of magnitude, so moving classification, extraction and formatting from the flagship to a cheap model is the single largest saving available. The remaining pits, by size: overfed context (the whole document every call), uncapped output (the model writes freely), failed requests billed anyway (timeouts and disconnects with no result), and the root one — not seeing where the money goes. Each has a concrete fix, ranked below from largest to smallest. The precondition: your platform must support multi-model switching and itemized ledgers — one OpenAI-compatible interface for Claude / GPT / Chinese models, and a console that shows every charge. With those two, cost reduction has handles.

1. Diagnose first: which category is your bill?

APIs bill per token, and both input (the context you send) and output (what the model generates) cost money. Before optimizing, spend ten minutes in the console ledger answering three questions: which model burns the most? Is input or output the bulk? Are failed requests being charged? The answers pick your first fix — most people find 'the flagship is 90% of spend, and half of those calls are simple work', which points to section 2.

2. Fix ①: model tiering — the largest saving

Top-tier models (Claude Opus / GPT flagships) and Chinese models (such as DeepSeek) differ in unit price by more than an order of magnitude, so moving half your simple calls to a cheap model can nearly halve the bill. The split: classification, extraction, summaries, formatting and first-draft translation — work 'with a right answer' — go to the cheap model; reasoning, complex code and user-facing critical output stay on the flagship. On one OpenAI-compatible interface the switch is a model string, so write the routing rule into code (pick the code by task type): quality holds, cost drops.

3. Fix ②: slimmer context — the neglected input side

In many scenarios input tokens outnumber output several times over: full history on every turn, the entire document on every request. Three actions: ① for chat, keep only the last few turns plus a summary instead of resending 50 turns; ② for long documents, retrieve or summarize first and feed only relevant passages, not the whole book; ③ audit the system prompt periodically — many grow bloated and half the rules are obsolete. Slimmer input saves on every call, multiplied by volume.

4. Fix ③: rein in output

Output tokens usually cost several times input, so 'make the model say less' is valuable. Two levers: set max_tokens — a sensible value per scenario (tens for classification, hundreds for summaries, looser for generation) to stop runaway output; constrain shape in the prompt — 'return JSON only', 'no more than three sentences', 'do not explain' — cheaper than truncating afterward because the surplus is never generated.

5. Fix ④: eliminate wasted tokens

Waste = a request fails midway on a timeout, disconnect or cutoff; you get nothing usable and are charged anyway. Negligible per call, insidious in batches — a job of tens of thousands of calls with a 5% failure rate is 5% pure waste. Remedies: a channel with stable latency, sensible timeouts and retries (with backoff, not a retry storm), and batches run in chunks with a failure list so only failures are rerun. The steadier the channel, the less waste — which is why 'stable' beats 'slightly cheaper' when choosing a relay.

6. Fix ⑤: reuse whatever you can

Do not pay twice for the same work: cache results locally for identical or near-identical requests (key on a hash of the request) and skip the call on a hit; fixed enumerations ('how do these 20 product categories map') belong in code, not in a prompt. Merge scattered requests into batches and run off-peak — cheaper and fewer rate-limit collisions. The impact depends on your repetition rate; high-repetition scenarios (support, moderation) often surprise.

7. Long term: a visible ledger keeps optimization going

Every fix above rests on one precondition: you can see where each dollar goes. When choosing a platform, confirm two things: usage detail by model and by key in the console, and a ledger that maps to individual calls. Then build a habit: review the spend ranking weekly, and the top burner is next week's optimization target. cocodot bills per token, exposes the full ledger in the console, takes Alipay, and one key covers Claude / GPT / Gemini / DeepSeek — tiered routing and reconciliation in one place.

8. Quick-start checklist

Copy as-is: ① sign up at cocodot (email verification grants a $0.5 trial credit), top up via Alipay, create a key; ② set base_url to https://cocodot.co/api/ai/v1; ③ high-frequency simple work on deepseek-v4-flash, hard work on mco-6 (Claude Opus) or mog-6 (GPT); ④ set max_tokens per task type; ⑤ after a week, review the ledger and rework the top burner per sections 2 and 3. Five steps, and most projects' bills step down noticeably.

Five fixes, ranked by savings

FixWaste targetedActionRough impact
① Model tieringExpensive models doing cheap workMove simple tasks to Chinese models; change the model codeLargest — unit prices differ 10x+
② Slimmer contextBloated input tokensSend only what is needed; summarize long text firstLarge — input is often the bulk
③ Capped outputModel writes without limitSet max_tokens; constrain length in the promptMedium
④ No wasted callsFailed requests billed anywayStable channel plus sensible timeouts and retriesMedium; hidden in batch jobs
⑤ Watch the ledgerNo idea where the money goesReview the console usage ranking weeklyPrecondition for the other four

FAQ

What is a sensible max_tokens?

By scenario: tens for classification / tagging; hundreds for summaries; generation as needed but always set. Principle: the normal upper bound for that scenario plus a little margin — it guards against runaway output, not normal output.

Are cheap Chinese models really good enough?

By scenario. For classification, extraction, formatting and summaries — work with a right answer — the gap to flagships is small while the price gap exceeds 10x; open-ended reasoning and complex code still warrant flagships. The right posture is routing, not replacing everything.

How much can caching save?

Depends on your repetition rate: support and moderation repeat a lot, so local result caching saves plenty; scenarios where every request differs gain little. Check the ledger for repeated patterns before building it.

How do I know whether I am wasting tokens?

Two numbers: failure rate (timeouts / disconnects) and charges on failed requests. If a batch job fails more than a few percent, fix the channel and timeout strategy before any other optimization — that is pure waste.

Do these methods require much code?

Mostly not: tiering is a model string, max_tokens is one parameter, slimmer context is a change to prompt assembly. The first three fit in an afternoon, and the bill usually moves the same week.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, $1 per active card per month; spending: $0.60 settlement fee on purchases under $20; a corresponding fee applies when the issuer charges one.

Service scope, pricing and limits →
Cutting AI API Costs: Model Tiering and Wasted Calls