Skip to content
Inference

Models and pricing

Token pricing

Prices are in USD per million tokens. Cached input is charged at the cached rate instead of the regular input rate.

Model IDInputCached inputOutput
deepseek-ai/DeepSeek-V4.1-Flash$0.12$0.0024$0.48
zai-org/GLM-5.3-Flash$0.12$0.024$0.40

Capabilities and limits

The models listed above support text and image input, streaming, reasoning, and function calling. Each supports a 1,048,576-token context, including output, with up to 131,072 output tokens per request. Reasoning tokens consume the output budget and are billed as output.

See the quickstart to make your first request.

Usage and billing

Token countChat Completions usageResponses usage
All input tokens, including cached inputprompt_tokensinput_tokens
Cached subset of inputprompt_tokens_details.cached_tokensinput_tokens_details.cached_tokens
All generated tokens, including reasoningcompletion_tokensoutput_tokens
Reasoning subset of outputcompletion_tokens_details.reasoning_tokensoutput_tokens_details.reasoning_tokens