Skip to main content
Multiple in-flight requests on the same API key run independently and in parallel. There is no head-of-line blocking, no per-key concurrency cap, and no implicit queuing — each request is dispatched to a backend as soon as it arrives. The only bound on a single key is the per-minute request ceiling described below. Inside that ceiling, fan out as wide as your workload needs.

Per-key rate limit

Each key carries a requests-per-minute (RPM) tier. When you exceed it the API returns 429 Too Many Requests with a Retry-After header — see Rate limits for the full headers. These are per-key, per-minute. They do not aggregate across keys on the same account today. The tier is stored on the key — request an upgrade from support@flex.ai if your steady-state load needs more.
We may introduce account-level concurrency or aggregate RPM caps as the platform scales. We will notify customers before tightening any existing per-key limit; we will not silently lower it.

Backpressure pattern

Bound your in-flight count to roughly RPM / 60 if you want to stay in a steady-state window, and retry on 429 with exponential backoff that respects Retry-After:
If you’re sustaining 429s after backoff, you’re above your tier’s steady-state capacity — request an upgrade rather than retrying harder. Tight retry loops don’t move you through the limit faster; they just burn your quota on rejected requests.

Ordering

Concurrent requests on one key are independent — completion order is not guaranteed to match submission order. If you need to correlate responses back to inputs, carry your own correlation id in the prompt or response metadata; don’t rely on arrival order. The id we return (chatcmpl-…) is unique per request and safe to use as a join key in your logs.