# Inference quickstart

Use your River API key with `https://api.river.ai/v1` and the OpenAI SDK or HTTP.

## Endpoints

| Endpoint | Use |
| --- | --- |
| `POST /v1/chat/completions` | Chat, images, tools, and streaming |
| `POST /v1/responses` | [Responses API](/guides/inference-controls/#responses-api) |
| `GET /v1/models` | List model IDs available through the API |

## Connect

Create an API key in the [River Console](https://console.river.ai/keys).
Set `RIVER_MODEL` to an ID from [Models and pricing](/guides/inference-models/):

```bash
pip install openai
export RIVER_API_KEY="rv_..."
export RIVER_MODEL="<model-id>"
```

Keep your API key on your server, outside browser code and source control.

## Stream a response

Save this as `inference.py` and run `python inference.py`:

```python
import os

from openai import OpenAI

with OpenAI(
    api_key=os.environ["RIVER_API_KEY"],
    base_url="https://api.river.ai/v1",
) as client:
    with client.chat.completions.create(
        model=os.environ["RIVER_MODEL"],
        messages=[{"role": "user", "content": "Say hello in one short sentence."}],
        max_tokens=1024,
        stream=True,
    ) as stream:
        for chunk in stream:
            if chunk.choices:
                print(chunk.choices[0].delta.content or "", end="", flush=True)
    print()
```

Send conversation history with each request. Unchanged prefixes can reuse
the cache when available.

## Call with curl

This returns a complete JSON response, with the answer in
`choices[0].message.content` and token counts in `usage`:

```bash
curl --fail-with-body https://api.river.ai/v1/chat/completions \
  -H "Authorization: Bearer $RIVER_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<JSON
  {
    "model": "$RIVER_MODEL",
    "messages": [{"role": "user", "content": "Say hello in one short sentence."}],
    "max_tokens": 1024
  }
JSON
```

For HTTP streaming, add `"stream": true` to the body and `-N` to curl.
The response uses server-sent events and ends with `data: [DONE]`.

For reasoning and structured output, see [API controls](/guides/inference-controls/).
For errors, retries, and timeouts, see [Limits and support](/guides/inference-limits/).
