Cursor and Cline With a Custom OpenAI Base URL: Setup, Verify and Fix Errors (2026)
Set a custom OpenAI base URL and key in Cursor, Cline and Roo Code. Exact fields, why override settings silently fail, a terminal test that isolates the problem, and error codes.
1. The three values that must all be right
A custom endpoint needs exactly three things, and a mistake in any one looks like 'cannot connect'. First, the base URL, ending at the /v1 level, for example https://cocodot.co/api/ai/v1. Some tools append /v1 themselves and some do not; when you see a 404, ask whether it was doubled or missing before anything else. Second, the model name: use the identifier the provider actually serves, copied from its model list, never remembered. Editors rarely list third-party models in their dropdown, so you will type it. Third, the key: copying from a console easily drags in a leading or trailing space or drops a character, so on any 401 the first move is to copy it again. These three failures have different causes and different fixes, so change one at a time.
2. Cursor: where the override lives, and the Verify trap
In Cursor's settings, open the models section and look for the custom OpenAI key option with an override base URL field. Enter the key and the base URL, then click Verify. This is the step most people skip, and Cursor does not apply the override until verification passes, which is why so many people conclude the override is not working. When verification fails, read the message: unauthorized points at the key, not found at the address or model name. Two further facts explain odd behaviour. Requests from your key are routed through Cursor's own servers, so an endpoint on localhost or behind your firewall cannot be reached; it has to be publicly accessible. And some Cursor features, such as built-in indexing and certain completions, use Cursor's own services regardless, so the override mainly affects chat and agent calls. Feature coverage with custom keys also changes between versions, so check the current behaviour if something is missing.
3. Cline and Roo Code
These VS Code extensions have an API Provider setting. Choose OpenAI Compatible, not OpenAI, and three fields appear: Base URL, API Key and Model ID. The Model ID must be typed by hand, because the list will not contain a third-party model. After saving, test in a new task. Old tasks can stay bound to the previous provider and model, which produces confusing results where the settings look right but traffic goes elsewhere. If the extension offers a way to set a custom context window or a maximum output limit, match it to the model you chose; too large a value gives context-length errors, too small a value truncates long edits mid-file.
4. Test from a terminal first
No error in the UI does not mean the connection works. The dependable check bypasses the editor: send one request with curl using the same base URL, key and model name. A JSON answer proves all three values are right, which splits the problem in half. If curl works but the editor fails, the fault is in the editor's configuration format: an extra slash, a model-name case mismatch, a stale session or a Verify step not clicked. If curl fails, the editor was never the problem. Verify the endpoint first, then debug the tool; doing it the other way round is how a ten-minute check becomes an afternoon.
curl https://cocodot.co/api/ai/v1/chat/completions \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"<model-id-from-the-list>","messages":[{"role":"user","content":"say ok"}]}'5. Error codes in plain language
401 or unauthorized means the key is wrong: copy it again, remove spaces, and make sure you pasted the key itself, not its label. 404 or model not found means a wrong model name, or /v1 missing or doubled in the address. 402 or insufficient means the balance is empty; top up. A timeout on a reasoning model is normal because those models think before they answer, so lengthen the tool's timeout; if ordinary models also time out, suspect the network path. 429 means rate limiting: reduce concurrency or retry after a short wait, and consider whether an agent is looping. A response that arrives but looks too short is usually a maximum-output setting rather than a provider fault.
6. Why people use a compatible endpoint at all
Two practical reasons. Payment: official subscriptions bill through overseas acquirers, and those score risk using the first 6 to 8 digits of the card number, the BIN, which tells them where the card was issued. A card issued locally in some regions is declined at a high rate, and a compatible endpoint that accepts your local payment method removes that step. Control: with a metered key, you see tokens and charges per call, you pick the model per tool, and you can run Cursor, Cline and a script from one balance. The trade-off is real. A subscription includes features that a raw key does not, so for some workflows the subscription remains the better tool; compare on a real week of work before moving.
7. Check the provider has not swapped the model
With any compatible endpoint, a reasonable worry is whether the model behind it is the one you pay for. You cannot tell by reading output. A workable method is a fixed set of probe prompts whose answers are characteristic of the real model, repeated over time. An open-source tool does this at probe.cocodot.co: give it any OpenAI-compatible base URL and a temporary key, and it works against any provider, including the one that publishes it. Run it when you first adopt an endpoint and again whenever behaviour suddenly changes. Use a throwaway key with a small balance, never your main one, and delete it afterwards.
8. A short debugging order that saves time
When a custom endpoint misbehaves, work in this order and stop at the first failure. One: curl the endpoint with the exact values. Two: confirm the model name against the provider's list, character for character. Three: check the base URL for a doubled or missing /v1. Four: in Cursor, click Verify and confirm the endpoint is public. Five: start a new session or task so old state does not interfere. Six: look at the balance and any per-key limits in the provider console. Seven: reduce the request, with a shorter prompt and no tools, to see if size or tool schemas are the trigger. Most cases end by step three.
Where each tool is configured and what trips people up
| Tool | Where | Common pitfall |
|---|---|---|
| Cursor | Settings, Models, then the custom OpenAI key and base URL fields | Override is ignored until you click Verify; endpoint must be public |
| Cline | Extension settings, API Provider: OpenAI Compatible | Model ID must be typed by hand; not in any dropdown |
| Roo Code | Same provider option as Cline | Test in a new task, since old tasks keep the old config |
| Continue | provider and apiBase in the config file | Reload the window after editing |
| Other tools | Look for base_url or API Base | Check whether the tool appends /v1 itself |