Inference
Models and pricing
Token pricing
Prices are in USD per million tokens. Cached input is charged at the cached rate instead of the regular input rate.
| Model ID | Input | Cached input | Output |
|---|---|---|---|
deepseek-ai/DeepSeek-V4.1-Flash | $0.12 | $0.0024 | $0.48 |
zai-org/GLM-5.3-Flash | $0.12 | $0.024 | $0.40 |
Capabilities and limits
The models listed above support text and image input, streaming, reasoning, and function calling. Each supports a 1,048,576-token context, including output, with up to 131,072 output tokens per request. Reasoning tokens consume the output budget and are billed as output.
See the quickstart to make your first request.
Usage and billing
| Token count | Chat Completions usage | Responses usage |
|---|---|---|
| All input tokens, including cached input | prompt_tokens | input_tokens |
| Cached subset of input | prompt_tokens_details.cached_tokens | input_tokens_details.cached_tokens |
| All generated tokens, including reasoning | completion_tokens | output_tokens |
| Reasoning subset of output | completion_tokens_details.reasoning_tokens | output_tokens_details.reasoning_tokens |