Cutting AI API Costs: Model Tiering, max_tokens, and Eliminating Wasted Calls (2026 Guide)
API bill creeping up? Five fixes ranked by savings: model tiering (biggest), slimmer context, capped output, no wasted calls, watching the ledger.
1. Diagnose first: which category is your bill?
APIs bill per token, and both input (the context you send) and output (what the model generates) cost money. Before optimizing, spend ten minutes in the console ledger answering three questions: which model burns the most? Is input or output the bulk? Are failed requests being charged? The answers pick your first fix — most people find 'the flagship is 90% of spend, and half of those calls are simple work', which points to section 2.
2. Fix ①: model tiering — the largest saving
Top-tier models (Claude Opus / GPT flagships) and Chinese models (such as DeepSeek) differ in unit price by more than an order of magnitude, so moving half your simple calls to a cheap model can nearly halve the bill. The split: classification, extraction, summaries, formatting and first-draft translation — work 'with a right answer' — go to the cheap model; reasoning, complex code and user-facing critical output stay on the flagship. On one OpenAI-compatible interface the switch is a model string, so write the routing rule into code (pick the code by task type): quality holds, cost drops.
3. Fix ②: slimmer context — the neglected input side
In many scenarios input tokens outnumber output several times over: full history on every turn, the entire document on every request. Three actions: ① for chat, keep only the last few turns plus a summary instead of resending 50 turns; ② for long documents, retrieve or summarize first and feed only relevant passages, not the whole book; ③ audit the system prompt periodically — many grow bloated and half the rules are obsolete. Slimmer input saves on every call, multiplied by volume.
4. Fix ③: rein in output
Output tokens usually cost several times input, so 'make the model say less' is valuable. Two levers: set max_tokens — a sensible value per scenario (tens for classification, hundreds for summaries, looser for generation) to stop runaway output; constrain shape in the prompt — 'return JSON only', 'no more than three sentences', 'do not explain' — cheaper than truncating afterward because the surplus is never generated.
5. Fix ④: eliminate wasted tokens
Waste = a request fails midway on a timeout, disconnect or cutoff; you get nothing usable and are charged anyway. Negligible per call, insidious in batches — a job of tens of thousands of calls with a 5% failure rate is 5% pure waste. Remedies: a channel with stable latency, sensible timeouts and retries (with backoff, not a retry storm), and batches run in chunks with a failure list so only failures are rerun. The steadier the channel, the less waste — which is why 'stable' beats 'slightly cheaper' when choosing a relay.
6. Fix ⑤: reuse whatever you can
Do not pay twice for the same work: cache results locally for identical or near-identical requests (key on a hash of the request) and skip the call on a hit; fixed enumerations ('how do these 20 product categories map') belong in code, not in a prompt. Merge scattered requests into batches and run off-peak — cheaper and fewer rate-limit collisions. The impact depends on your repetition rate; high-repetition scenarios (support, moderation) often surprise.
7. Long term: a visible ledger keeps optimization going
Every fix above rests on one precondition: you can see where each dollar goes. When choosing a platform, confirm two things: usage detail by model and by key in the console, and a ledger that maps to individual calls. Then build a habit: review the spend ranking weekly, and the top burner is next week's optimization target. cocodot bills per token, exposes the full ledger in the console, takes Alipay, and one key covers Claude / GPT / Gemini / DeepSeek — tiered routing and reconciliation in one place.
8. Quick-start checklist
Copy as-is: ① sign up at cocodot (email verification grants a $0.5 trial credit), top up via Alipay, create a key; ② set base_url to https://cocodot.co/api/ai/v1; ③ high-frequency simple work on deepseek-v4-flash, hard work on mco-6 (Claude Opus) or mog-6 (GPT); ④ set max_tokens per task type; ⑤ after a week, review the ledger and rework the top burner per sections 2 and 3. Five steps, and most projects' bills step down noticeably.
Five fixes, ranked by savings
| Fix | Waste targeted | Action | Rough impact |
|---|---|---|---|
| ① Model tiering | Expensive models doing cheap work | Move simple tasks to Chinese models; change the model code | Largest — unit prices differ 10x+ |
| ② Slimmer context | Bloated input tokens | Send only what is needed; summarize long text first | Large — input is often the bulk |
| ③ Capped output | Model writes without limit | Set max_tokens; constrain length in the prompt | Medium |
| ④ No wasted calls | Failed requests billed anyway | Stable channel plus sensible timeouts and retries | Medium; hidden in batch jobs |
| ⑤ Watch the ledger | No idea where the money goes | Review the console usage ranking weekly | Precondition for the other four |