Skip to main content
Inference

Token Factory changelog

Changes to the Token Factory are now published here — models arriving and leaving, pricing changes, and changes to the API.For what is serving right now, GET /v1/models remains the authoritative list and carries current per-token pricing for every model on it.
Inference

New models: GLM 5.3 Flash and Qwen3.8 27B

Both are serving now on the Inference API, and both read images as well as text.
  • GLM 5.3 Flash — $0.075 per million input tokens, $0.25 per million output, with a 1M-token context window.
  • Qwen3.8 27B — $0.42 per million input tokens, $3.00 per million output, with a 262K-token context window.
Call either by its id through the usual endpoint: