Skip to content
Inference

Inference quickstart

Use your River API key with https://api.river.ai/v1 and the OpenAI SDK or HTTP.

Endpoints

EndpointUse
POST /v1/chat/completionsChat, images, tools, and streaming
POST /v1/responsesResponses API
GET /v1/modelsList model IDs available through the API

Connect

Create an API key in the River Console. Set RIVER_MODEL to an ID from Models and pricing:

pip install openai
export RIVER_API_KEY="rv_..."
export RIVER_MODEL="<model-id>"

Keep your API key on your server, outside browser code and source control.

Stream a response

Save this as inference.py and run python inference.py:

import os

from openai import OpenAI

with OpenAI(
    api_key=os.environ["RIVER_API_KEY"],
    base_url="https://api.river.ai/v1",
) as client:
    with client.chat.completions.create(
        model=os.environ["RIVER_MODEL"],
        messages=[{"role": "user", "content": "Say hello in one short sentence."}],
        max_tokens=1024,
        stream=True,
    ) as stream:
        for chunk in stream:
            if chunk.choices:
                print(chunk.choices[0].delta.content or "", end="", flush=True)
    print()

Send conversation history with each request. Unchanged prefixes can reuse the cache when available.

Call with curl

This returns a complete JSON response, with the answer in choices[0].message.content and token counts in usage:

curl --fail-with-body https://api.river.ai/v1/chat/completions \
  -H "Authorization: Bearer $RIVER_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<JSON
  {
    "model": "$RIVER_MODEL",
    "messages": [{"role": "user", "content": "Say hello in one short sentence."}],
    "max_tokens": 1024
  }
JSON

For HTTP streaming, add "stream": true to the body and -N to curl. The response uses server-sent events and ends with data: [DONE].

For reasoning and structured output, see API controls. For errors, retries, and timeouts, see Limits and support.