model, messages, temperature, top_p, stop, seed, user | Yes | A sampling field you omit may be filled with the model’s recommended default (for example top_p: 0.95 on DeepSeek-V4-Flash-0731, the value its model card recommends for agentic use). Only omitted fields are filled: an explicit value, including temperature: 0, is always sent as-is. The response names what was filled in x-flexai-sampling-defaulted (for example top_p); no header means nothing was filled. The same applies on /v1/responses. |
max_tokens | Yes | An explicit value is honored as-is (max_completion_tokens is honored too, though not yet listed in supported_parameters). Non-streaming requests that set neither are bounded by a ~95 s generation budget — stream for longer outputs — and, on non-reasoning models only, default to 2048 output tokens; the response then carries x-flexai-max-tokens-defaulted: 2048. Reasoning models (always-thinking models such as GLM-5.3-Flash, gpt-oss-120b, Step-3.7-Flash, and the hybrids) get no default, because the reasoning would spend it and return an empty answer. A non-streamed response that ends with finish_reason: "length" and no visible content or tool call carries x-flexai-empty-at-length: true — the model was cut off, not incapable; retry with a larger max_tokens or stream. On tool turns set 4096 or more: a tool call cut off by max_tokens returns finish_reason: "length" with no tool_calls. |
stream, stream_options.include_usage | Yes | Set include_usage: true when streaming for correct billing. A stream the gateway ends because the model was repeating itself carries finish_reason: "length" and flexai_stopped: "degeneration" on its terminal chunks ("ceiling" when it was cut at the token ceiling while repeating) — see Runaway streams. |
tools, tool_choice | Yes | On models labeled tool_use in the catalog. |
response_format: { type: "json_object" } | Partial | On models listing response_format in their /v1/models supported_parameters; other models return a 400 with param: "response_format". |
response_format: { type: "json_schema", ... } | Partial | On models listing structured_outputs in their /v1/models supported_parameters; other models return a 400 with param: "response_format" (there is no automatic fallback to json_object). |
logprobs, top_logprobs | No | Not at launch. |
n (multiple completions per request) | No | Gateway returns a single choice; sending n > 1 returns a 400 with param: "n". |
presence_penalty, frequency_penalty | Yes | — |
reasoning_effort | Yes | Accepted on every chat model; never a 400. Levels: minimal, low, medium, high, max (case-insensitive). Each model’s chat template understands its own words, so the gateway translates per model: gpt-oss-120b / gpt-oss-20b take low, medium, high natively; DeepSeek-V4-Flash-0731 — low/minimal keep reasoning off, medium/high select its high level, max its maximum; GLM-5.2 — minimal/low off, medium/high → high, max → max; Qwen3.8-27B — minimal off, low/medium as sent, high/max → xhigh; GLM-5.3-Flash (always-thinking) — minimal/low → low, medium/high → high, max → max; Muse-Glimmer-30B (always-thinking) — minimal/low → low, medium → medium, high/max → high. These maps are the gateway’s and apply whether or not supported_parameters lists reasoning_effort for the model. On a model with no depth dial (including Step-3.7-Flash, whose template does not act on the word) the parameter is removed and the response carries x-flexai-dropped-params: reasoning_effort, so a silent drop is never mistaken for a setting that took. Reasoning is off by default on every model with an on/off switch; a reasoning_effort above low turns it on where the model has a dial, thinking (below) everywhere. Reasoning text is returned in message.reasoning_content (delta.reasoning_content when streaming) and counted in usage.completion_tokens_details.reasoning_tokens. That count is the serving engine’s own token count for the reasoning (on streams served by the new gateway it replaces an approximate recount); it is informational, since reasoning tokens are already billed inside completion_tokens. |
thinking | Yes | {"type": "enabled"} or {"type": "disabled"} — the reasoning on/off switch on reasoning-capable models: DeepSeek-V4-Flash-0731 and the hybrid-thinking models (gemma-4-31b-it, gemma-4-26B-A4B-it, Qwen3.5-9B, Qwen3.6-35B-A3B-FP8, GLM-5.2, Qwen3.8-27B, GLM-4.5-Air-FP8, Qwen3-8B-FP8, NVIDIA-Nemotron-3.5-Lightning-30B-A3B, which are off by default; where the model also has a reasoning_effort dial — GLM-5.2, Qwen3.8-27B — an effort above low turns it on too, otherwise {"type": "enabled"} is the only way). It takes precedence over reasoning_effort: disabled turns reasoning off even if reasoning_effort is also sent; enabled alone turns it on at the model’s default depth, and combined with a reasoning_effort uses that depth (low/minimal floor to high). On a model that cannot stop reasoning — any model whose catalog row declares it always reasons (a name containing Thinking is the usual hint; the set moves with the fleet, so test it rather than keep a list), and gpt-oss-120b / gpt-oss-20b, whose reasoning_effort dial bottoms out at low — {"type": "disabled"} is a 400 naming thinking; the gpt-oss message points at reasoning_effort: "low" as the least reasoning it can do. Omit the parameter to accept the model’s default. On a model with no reasoning switch (Llama-3.3-70B-Instruct-FP8, Meta-Llama-3.1-8B-Instruct-FP8, Mistral-Nemo-Instruct-2407-FP8, Qwen3-Coder-30B-A3B-Instruct-FP8) a well-formed thinking is removed and the response carries x-flexai-dropped-params: thinking (reasoning when sent in the OpenRouter dialect); a malformed value is a 400 on every model. The raw chat_template_kwargs.enable_thinking / chat_template_kwargs.thinking switches must be a boolean (or null, meaning unset) on every model, or the request is a 400 naming the field. Not yet listed in supported_parameters. See tool use. |