Per-key rate limit
Each key carries a requests-per-minute (RPM) tier. When you exceed it the API returns429 Too Many Requests with a Retry-After header — see Rate limits for the full headers.
These are per-key, per-minute. They do not aggregate across keys on the same account today. The tier is stored on the key — request an upgrade from support@flex.ai if your steady-state load needs more.
We may introduce account-level concurrency or aggregate RPM caps as the platform scales. We will notify customers before tightening any existing per-key limit; we will not silently lower it.
Backpressure pattern
Bound your in-flight count to roughlyRPM / 60 if you want to stay in a steady-state window, and retry on 429 with exponential backoff that respects Retry-After:
Ordering
Concurrent requests on one key are independent — completion order is not guaranteed to match submission order. If you need to correlate responses back to inputs, carry your own correlation id in the prompt or response metadata; don’t rely on arrival order. Theid we return (chatcmpl-…) is unique per request and safe to use as a join key in your logs.
Related
- Batching — batch processing is coming soon.
- Billing & quotas — tier limits and rate-limit response headers.
- Errors — the full 429 body.