Skip to content
Inference

API controls

Reasoning

ModelChat: reasoning_effortResponses: reasoning.effort
DeepSeek-V4.1-Flashnone (off), high, maxnone, high, max
GLM-5.3-Flashlow, high, maxlow, high, max

GLM requires reasoning. Chat returns reasoning in reasoning_content and the answer in content (message for JSON, delta for streaming). Both count toward the output budget and price.

Sampling

SettingChat CompletionsResponses
Output budgetmax_tokensmax_output_tokens
Temperaturetemperature: 0–2temperature: 0–2
Nucleus samplingtop_p: 0–10 < top_p <= 1
Streamingstream: truestream: true

Chat also accepts stop as a string or list of strings. If finish_reason is length, increase the output budget within the context limit. See model limits and tool calling.

Structured output

JSON Schema is supported on GLM; it is not advertised for DeepSeek. Add this to a GLM Chat Completions request:

{
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "result",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {"ok": {"type": "boolean"}},
        "required": ["ok"],
        "additionalProperties": false
      }
    }
  }
}

For Responses, use text.format with type, name, strict, and schema. Validate the returned JSON and check for truncation.

Responses API

Inside the quickstart's client block:

response = client.responses.create(
    model=os.environ["RIVER_MODEL"],
    input="What is 2 + 2?",
    reasoning={"effort": "high"},
    max_output_tokens=1024,
    store=False,
)
print(response.output_text)

Send full history in input. store: true, background: true, previous_response_id, and conversation are unsupported. Successful streams end with response.completed.