cocodot
← Back to guides
Local card declined for overseas AI? cocodot: one card + one key
ExplainersUpdated 2026-07

429 Too Many Requests From an LLM API: Rate Limits, Concurrency, and Retries That Work

429 means you exceeded a quota, but there are three different quotas and they fail differently. Identify which one you hit before changing anything, then fix it in the order that works.

TL;DR: 429 Too Many Requests means you exceeded a quota — but there are three, and the fix differs. RPM (requests per minute), TPM (tokens per minute, the one long prompts blow through), and concurrency (how many are in flight at once). Read the error body: most providers name which limit you hit. Then, in this order: (1) exponential backoff with jitter — on 429, back off and retry; hammering harder makes the queue worse, and without jitter your retries synchronize into a new spike. (2) Client-side rate limiting — a semaphore that turns fire all 1,000 at once into a controlled 5 at a time; mandatory for batch jobs. (3) Queue and smooth so bursts get spread out. The most common misdiagnosis: you believe concurrency is 10, but each worker runs its own limiter, and 10 workers × 10 each is 100.

Read the error before changing anything

A 429 body usually names the limit that was hit and often includes a retry-after hint. Skipping this step is why people fix the wrong thing — halving concurrency does nothing if you were hitting TPM with a single enormous prompt. Log the full error body on 429, not just the status code; the difference between the three cases is right there and takes seconds to read.

Exponential backoff, with jitter

On a 429, wait and retry with the delay doubling each attempt, capped at a ceiling, with a hard limit on attempts. The jitter matters more than people expect: without random variation, every client that got throttled at the same instant retries at the same instant, recreating the spike that caused the problem. Add a random fraction to each delay. Also respect retry-after when the provider sends it — that is the provider telling you exactly how long to wait.

Rate limit on your side, not just theirs

Backoff is reactive; a client-side limiter is preventive and it is what keeps batch jobs from melting. Wrap calls in a semaphore sized to your actual concurrency allowance, and if you are hitting RPM, add a token-bucket limiter as well. The rewrite is small — the loop that fires everything becomes a loop that acquires a slot first — and it converts a job that fails halfway into one that simply runs a bit slower.

Queue and smooth the bursts

If your traffic arrives in bursts — a nightly job, a user action that fans out — a queue between the trigger and the API turns a spike into a steady stream. This is more work than the first two fixes, so do it when you have burst-shaped traffic rather than as a default. For steady traffic, a limiter is enough.

The mistake almost everyone makes once

Concurrency limits are global, and limiters usually are not. Ten workers each holding a semaphore of ten gives you a hundred in flight while your config says ten. If you run multiple processes, the limiter has to be shared — a central queue, or a limit divided across workers. When the numbers look right and you still get 429, this is the first place to look.

Your gateway has limits too

If you call through a gateway rather than the provider directly, there are two limiters in the path: theirs upstream and the gateway's own. A gateway that pools capacity across customers can absorb bursts an individual account could not — or can become the bottleneck itself. Worth asking any provider what their concurrency allowance is before you build around an assumption. cocodot does not impose an additional per-key concurrency cap beyond what upstream enforces, so what you see is the provider's own limit.

Checklist when you are stuck

1. Log the full 429 body and identify which limit. 2. Confirm your effective concurrency by counting in-flight requests, not by reading config. 3. Verify backoff has jitter and a cap on attempts. 4. For TPM, measure actual tokens per request — a context that grew over time is a common silent cause. 5. If none of that explains it, check whether another process shares the same key.

A 429 that is not a rate limit: the empty balance

Some providers return 429 for exhausted credit or quota, not for speed. OpenAI's API, for example, uses 429 with an insufficient-quota code when the account has no balance. Retrying such an error is pure waste: it will fail identically forever. The body tells you — a rate-limit 429 mentions requests or tokens per minute, while a quota 429 mentions billing, credit or plan. Build the distinction into your retry logic: retry the first kind with backoff, and stop and alert on the second.

max_tokens counts before you use it

On several providers, the tokens-per-minute figure is estimated at request time from your prompt plus the `max_tokens` you allowed, not from what the model actually generates. A request that sets a huge `max_tokens` 'just in case' reserves that capacity and can trigger TPM throttling even when the real outputs are short. Set `max_tokens` to what the task needs, and watch whether 429s drop. It is one of the few fixes that costs nothing and needs no new code.

Long streams hold their slot

A concurrency slot stays occupied until the response finishes, and streaming responses can run for a minute or more. Ten long agent turns can pin ten slots for their whole duration while your request rate looks tiny. If 429s appear with low RPM and modest tokens, count in-flight requests over time rather than requests per minute — the limit you are hitting is probably concurrency, and the cure is capping workers or shortening outputs.

Three limits, three symptoms

LimitWhat it countsTypical triggerFirst fix
RPMRequests per minuteMany small callsClient-side rate limiter
TPMTokens per minuteLong prompts or big contextBatch fewer, trim context
ConcurrencyIn-flight requestsParallel workersSemaphore across all workers

FAQ

Should I just retry immediately on 429?

No. Immediate retries add load to the exact resource that is already saturated and usually extend the outage. Back off, add jitter, cap the attempts.

Does a bigger balance raise my rate limits?

Sometimes with direct provider accounts, where tiers scale with usage history. It is not a general rule, and it never fixes a client-side concurrency bug.

How do I avoid 429 Too Many Requests altogether?

You cannot promise zero, but you can make them rare: a shared client-side limiter sized to your real allowance, `max_tokens` set to what the task needs, backoff with jitter for the occasional miss, and a queue in front of bursty jobs.

Why do I get 429 on the OpenAI API with low usage?

Two common reasons: your balance or quota is exhausted (a billing 429, not a speed limit), or `max_tokens` is set so high that the estimated token load hits the per-minute cap. Read the error body to see which.

What is the difference between 429 and 503 or 529?

429 means you exceeded your allowance; retrying later on your schedule is correct. 503 and similar overload codes mean the provider is short of capacity regardless of your usage — backoff still helps, but reducing your own rate will not.

About cocodot

cocodot is a payment and AI access service for developers and cross-border teams in mainland China. It provides US-BIN virtual cards issued by a licensed institution — used to pay for overseas subscriptions and ad accounts — and an OpenAI-compatible AI API gateway for calling Claude, GPT and Gemini from within mainland China. Both share one wallet, funded by Alipay and accounted in USD. Card: $9.9 to open, 3% to load, $1 per active card per month; spending: $0.60 settlement fee on purchases under $20; a corresponding fee applies when the issuer charges one.

Service scope, pricing and limits →
429 Too Many Requests on an LLM API: Fixes That Work