# Models and pricing

## Token pricing

Prices are in **USD per million tokens**. Cached input is charged at the cached
rate instead of the regular input rate.

| Model ID | Input | Cached input | Output |
| --- | ---: | ---: | ---: |
| `deepseek-ai/DeepSeek-V4.1-Flash` | $0.12 | $0.0024 | $0.48 |
| `zai-org/GLM-5.3-Flash` | $0.12 | $0.024 | $0.40 |

## Capabilities and limits

The models listed above support text and image input, streaming, reasoning, and function
calling. Each supports a **1,048,576-token context**, including output, with
up to **131,072 output tokens** per request. Reasoning tokens consume the
output budget and are billed as output.

See the [quickstart](/guides/inference/) to make your first request.

## Usage and billing

| Token count | Chat Completions `usage` | Responses `usage` |
| --- | --- | --- |
| All input tokens, including cached input | `prompt_tokens` | `input_tokens` |
| Cached subset of input | `prompt_tokens_details.cached_tokens` | `input_tokens_details.cached_tokens` |
| All generated tokens, including reasoning | `completion_tokens` | `output_tokens` |
| Reasoning subset of output | `completion_tokens_details.reasoning_tokens` | `output_tokens_details.reasoning_tokens` |
