429 Too Many Requests From an LLM API: Rate Limits, Concurrency, and Retries That Work
429 means you exceeded a quota, but there are three different quotas and they fail differently. Identify which one you hit before changing anything, then fix it in the order that works.
Read the error before changing anything
A 429 body usually names the limit that was hit and often includes a retry-after hint. Skipping this step is why people fix the wrong thing — halving concurrency does nothing if you were hitting TPM with a single enormous prompt. Log the full error body on 429, not just the status code; the difference between the three cases is right there and takes seconds to read.
Exponential backoff, with jitter
On a 429, wait and retry with the delay doubling each attempt, capped at a ceiling, with a hard limit on attempts. The jitter matters more than people expect: without random variation, every client that got throttled at the same instant retries at the same instant, recreating the spike that caused the problem. Add a random fraction to each delay. Also respect retry-after when the provider sends it — that is the provider telling you exactly how long to wait.
Rate limit on your side, not just theirs
Backoff is reactive; a client-side limiter is preventive and it is what keeps batch jobs from melting. Wrap calls in a semaphore sized to your actual concurrency allowance, and if you are hitting RPM, add a token-bucket limiter as well. The rewrite is small — the loop that fires everything becomes a loop that acquires a slot first — and it converts a job that fails halfway into one that simply runs a bit slower.
Queue and smooth the bursts
If your traffic arrives in bursts — a nightly job, a user action that fans out — a queue between the trigger and the API turns a spike into a steady stream. This is more work than the first two fixes, so do it when you have burst-shaped traffic rather than as a default. For steady traffic, a limiter is enough.
The mistake almost everyone makes once
Concurrency limits are global, and limiters usually are not. Ten workers each holding a semaphore of ten gives you a hundred in flight while your config says ten. If you run multiple processes, the limiter has to be shared — a central queue, or a limit divided across workers. When the numbers look right and you still get 429, this is the first place to look.
Your gateway has limits too
If you call through a gateway rather than the provider directly, there are two limiters in the path: theirs upstream and the gateway's own. A gateway that pools capacity across customers can absorb bursts an individual account could not — or can become the bottleneck itself. Worth asking any provider what their concurrency allowance is before you build around an assumption. cocodot does not impose an additional per-key concurrency cap beyond what upstream enforces, so what you see is the provider's own limit.
Checklist when you are stuck
1. Log the full 429 body and identify which limit. 2. Confirm your effective concurrency by counting in-flight requests, not by reading config. 3. Verify backoff has jitter and a cap on attempts. 4. For TPM, measure actual tokens per request — a context that grew over time is a common silent cause. 5. If none of that explains it, check whether another process shares the same key.
A 429 that is not a rate limit: the empty balance
Some providers return 429 for exhausted credit or quota, not for speed. OpenAI's API, for example, uses 429 with an insufficient-quota code when the account has no balance. Retrying such an error is pure waste: it will fail identically forever. The body tells you — a rate-limit 429 mentions requests or tokens per minute, while a quota 429 mentions billing, credit or plan. Build the distinction into your retry logic: retry the first kind with backoff, and stop and alert on the second.
max_tokens counts before you use it
On several providers, the tokens-per-minute figure is estimated at request time from your prompt plus the `max_tokens` you allowed, not from what the model actually generates. A request that sets a huge `max_tokens` 'just in case' reserves that capacity and can trigger TPM throttling even when the real outputs are short. Set `max_tokens` to what the task needs, and watch whether 429s drop. It is one of the few fixes that costs nothing and needs no new code.
Long streams hold their slot
A concurrency slot stays occupied until the response finishes, and streaming responses can run for a minute or more. Ten long agent turns can pin ten slots for their whole duration while your request rate looks tiny. If 429s appear with low RPM and modest tokens, count in-flight requests over time rather than requests per minute — the limit you are hitting is probably concurrency, and the cure is capping workers or shortening outputs.
Three limits, three symptoms
| Limit | What it counts | Typical trigger | First fix |
|---|---|---|---|
| RPM | Requests per minute | Many small calls | Client-side rate limiter |
| TPM | Tokens per minute | Long prompts or big context | Batch fewer, trim context |
| Concurrency | In-flight requests | Parallel workers | Semaphore across all workers |