Inference quickstart
Use your River API key with https://api.river.ai/v1 and the OpenAI SDK or HTTP.
Endpoints
| Endpoint | Use |
|---|---|
POST /v1/chat/completions | Chat, images, tools, and streaming |
POST /v1/responses | Responses API |
GET /v1/models | List model IDs available through the API |
Connect
Create an API key in the River Console.
Set RIVER_MODEL to an ID from Models and pricing:
pip install openai
export RIVER_API_KEY="rv_..."
export RIVER_MODEL="<model-id>"Keep your API key on your server, outside browser code and source control.
Stream a response
Save this as inference.py and run python inference.py:
import os
from openai import OpenAI
with OpenAI(
api_key=os.environ["RIVER_API_KEY"],
base_url="https://api.river.ai/v1",
) as client:
with client.chat.completions.create(
model=os.environ["RIVER_MODEL"],
messages=[{"role": "user", "content": "Say hello in one short sentence."}],
max_tokens=1024,
stream=True,
) as stream:
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()Send conversation history with each request. Unchanged prefixes can reuse the cache when available.
Call with curl
This returns a complete JSON response, with the answer in
choices[0].message.content and token counts in usage:
curl --fail-with-body https://api.river.ai/v1/chat/completions \
-H "Authorization: Bearer $RIVER_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<JSON
{
"model": "$RIVER_MODEL",
"messages": [{"role": "user", "content": "Say hello in one short sentence."}],
"max_tokens": 1024
}
JSONFor HTTP streaming, add "stream": true to the body and -N to curl.
The response uses server-sent events and ends with data: [DONE].
For reasoning and structured output, see API controls. For errors, retries, and timeouts, see Limits and support.