> ## Documentation Index
> Fetch the complete documentation index at: https://docs.flex.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Regions & Data Residency

> The US and EU API endpoints, how a workspace is pinned to one, and what that changes about models, keys, and billing.

FlexAI serves the Inference API from two regions. Which one you use is decided once, when your workspace is created, and everything else on this page follows from that single choice.

| Region | API base URL | What it is |
| - | - | - |
| `us` (default) | `https://api.flex.ai/v1` | The global endpoint. Served from our worldwide capacity. |
| `eu` | `https://eu.api.flex.ai/v1` | The EU endpoint. Your prompts, completions, keys and usage records stay inside the EU. |

If you have never been asked about a region, you are in `us` and `https://api.flex.ai/v1` is your endpoint.

## What a region actually is

A region is a physical fact about where a workspace runs, not a label on its rows. Each region is a **separate stack** with its own database, gateway and GPU fleet, and a workspace's API keys, request content and usage records live in exactly one of them.

The two differ in how far their serving reaches:

* **The EU endpoint is strict.** It routes only to GPU capacity inside the EU, so an EU workspace's prompts and completions are processed in the EU. That is enforced at the routing layer — the stack will not list capacity from anywhere else — rather than checked per request.
* **The US endpoint is global.** Inference is served from our capacity wherever it sits, which includes EU hardware. It is the right default for performance and model coverage, and it is **not** a commitment that a given request is processed in any particular country.

### What the EU guarantee covers, and what it does not

The boundary is a real one, and it is better learned here than in a procurement questionnaire.

**Inside the EU stack:** your API keys, the content of your requests and responses, the authoritative records of your usage and spend, and the inference itself. For chat, completions and embeddings, request and response bodies are not retained once the request is served. Video generation is the exception, because it cannot work otherwise: the prompt and the generated clip are stored, inside the EU stack, until a cleanup job reaps them once the generation has finished, within the 24 to 48 hour window set out in our [Privacy Policy](https://flex.ai/privacy-policy).

**Held centrally, in the US:** everything in the account layer above the workspace. That layer is a single global control plane rather than something duplicated per region, so treat the rule as the whole layer rather than a list to check against — it covers, among other things, your sign-in identity and sessions, your organization and its membership list, the billing account with its payment methods and ledger, and the records of any identity verification you complete, which includes government-ID checks run by our verification provider. It is also why a single dashboard manages your workspaces in either region, reachable at [platform.flex.ai](https://platform.flex.ai). The EU endpoint carries the API and the dashboard's own calls against it, nothing else; opening `eu.api.flex.ai` in a browser returns 404, which is correct rather than broken.

**Operational telemetry is a separate question.** The metrics, traces and product analytics we use to run and monitor the service are not part of the regional boundary described above, and some of them carry usage-derived figures. They do not carry your prompts or completions. If your requirement covers that category too, raise it with us explicitly rather than reading it off this page.

So the EU guarantee is about where your **inference and its data** are processed and stored, and that is the part most residency requirements are written about. If your requirement extends to account and billing metadata as well, talk to us at [support@flex.ai](mailto:support@flex.ai) before you build — do not infer it from this page.

## Choosing a region

You pick the region when you create a workspace, and **the choice is permanent**. There is no migration path and no setting that changes it afterwards — moving a workspace between regions would mean moving the data the region exists to contain, which is the one thing the design rules out.

An organization can hold **one workspace per region**, so if you need both, create a second workspace rather than trying to move the first. The two are independent: separate keys, separate balances, separate usage.

If a workspace ended up in the wrong region, it cannot be moved — not from the dashboard and not by us. Contact [support@flex.ai](mailto:support@flex.ai) and we will help you set up a new workspace in the right one.

## Keys are region-scoped

An API key belongs to the workspace that minted it, and therefore to that workspace's region. The two stacks hold separate key stores, so an EU key does not exist as far as `https://api.flex.ai/v1` is concerned, and a US key does not exist to `https://eu.api.flex.ai/v1`.

That means **a key presented to the wrong host fails authentication** — `401`, the same answer a revoked or mistyped key gets — rather than returning anything that mentions regions. If a key that works in one place is suddenly unauthorized in another, check the host before you suspect the key.

The fix is always to call the host for your own region, never to re-mint the key.

(There is a separate `403` carrying the code `region_not_served`, which the [errors reference](/inference-api/reference/errors#region-not-served) documents. You are very unlikely to meet it: it guards against a stack somehow holding a workspace from a region it does not serve, which is an operational fault on our side rather than something a wrong base URL produces.)

## What each region serves

The EU endpoint serves less than the US one, for two independent reasons:

* **Capacity.** The EU endpoint routes only to EU hardware, so a model with no EU capacity behind it is not served there. The reverse is possible too — the US endpoint can only route to another region's hardware when the connection to it resolves — so treat the two catalogs as *different*, not as one containing the other.
* **Licensing.** Some models carry license or provenance terms we are not prepared to serve into the EU, and those are withheld there regardless of capacity.

Neither of these is a list you can read ahead of time, and both move as the fleet and the legal picture change. So treat `GET /v1/models` **against the host you are calling** as the authoritative answer, exactly as the [model discovery guide](/inference-api/guides/model-discovery) describes:

```bash theme={null}
curl https://eu.api.flex.ai/v1/models \
  -H "Authorization: Bearer $FLEXAI_API_KEY"
```

A model your region does not serve is absent from that list, and calling it anyway returns a `404` — but which `404` depends on why, and the difference is worth knowing when you are debugging:

* A model with **no capacity** in your region answers `404` with the code `model_deactivated`. The id is one we recognize; there is nothing behind it here right now.
* A model **not cleared for your region** answers the ordinary unknown-model `404` instead, with no separate code saying so. Usually that means the licensing review for it has not been completed yet rather than that a problem was found. We do not publish a per-region list, so treat absence from `/v1/models` as the answer.

If you are porting code from a US workspace, re-check your model ids against the EU list rather than assuming they carry over.

## Billing currency

A workspace's billing follows its region. US accounts are charged and invoiced in **US dollars**; EU accounts are charged and invoiced in **euros**, with EU VAT handled on the invoice.

Underneath that, prices stay in dollars everywhere. The model catalog, your usage figures and your wallet balance are all USD, and a euro top-up is converted to dollars exactly once — at the moment the card is charged, at the European Central Bank reference rate fixed for that calendar month, recorded against the payment. Nothing is re-converted afterwards, so the dollar value of credit you have already bought never moves with the exchange rate.

The practical upshot: you type euros, Stripe charges euros, your invoice reads euros, and everything you see about *usage* reads dollars. Top-up minimums and presets are set natively per currency rather than converted, so the euro figures are round numbers too.

See [Billing & Quotas](/inference-api/reference/billing) for how metering, per-key rate limits and the `402` response work. Those behave the same in both regions; the edge's own flood limit and some billing mechanics differ, so read a limit off the response headers rather than assuming it carries across.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.