Inference
Token Factory changelog
Changes to the Token Factory are now published here — models arriving and leaving, pricing changes, and changes to the API.For what is serving right now,GET /v1/models remains the authoritative list and carries current per-token pricing for every model on it.Inference
New models: GLM 5.3 Flash and Qwen3.8 27B
Both are serving now on the Inference API, and both read images as well as text.- GLM 5.3 Flash — $0.075 per million input tokens, $0.25 per million output, with a 1M-token context window.
- Qwen3.8 27B — $0.42 per million input tokens, $3.00 per million output, with a 262K-token context window.