# API controls

## Reasoning

| Model | Chat: `reasoning_effort` | Responses: `reasoning.effort` |
| --- | --- | --- |
| DeepSeek-V4.1-Flash | `none` (off), `high`, `max` | `none`, `high`, `max` |
| GLM-5.3-Flash | `low`, `high`, `max` | `low`, `high`, `max` |

GLM requires reasoning. Chat returns reasoning in
`reasoning_content` and the answer in `content` (`message` for JSON, `delta`
for streaming). Both count toward the output budget and price.

## Sampling

| Setting | Chat Completions | Responses |
| --- | --- | --- |
| Output budget | `max_tokens` | `max_output_tokens` |
| Temperature | `temperature`: `0`–`2` | `temperature`: `0`–`2` |
| Nucleus sampling | `top_p`: `0`–`1` | `0 < top_p <= 1` |
| Streaming | `stream: true` | `stream: true` |

Chat also accepts `stop` as a string or list of strings. If `finish_reason` is
`length`, increase the output budget within the context limit.
See [model limits](/guides/inference-models/#capabilities-and-limits) and
[tool calling](/guides/inference-tools/).

## Structured output

JSON Schema is supported on GLM; it is not advertised for DeepSeek.
Add this to a GLM Chat Completions request:

```json
{
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "result",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {"ok": {"type": "boolean"}},
        "required": ["ok"],
        "additionalProperties": false
      }
    }
  }
}
```

For Responses, use `text.format` with `type`, `name`, `strict`, and `schema`.
Validate the returned JSON and check for truncation.

## Responses API

Inside the [quickstart's client block](/guides/inference/):

```python
response = client.responses.create(
    model=os.environ["RIVER_MODEL"],
    input="What is 2 + 2?",
    reasoning={"effort": "high"},
    max_output_tokens=1024,
    store=False,
)
print(response.output_text)
```

Send full history in `input`. `store: true`, `background: true`,
`previous_response_id`, and `conversation` are unsupported.
Successful streams end with `response.completed`.
