Inference
Limits and support
Limits
| Setting | Limit |
|---|---|
| Context and output | Per-model limits |
| Request body | 50 MiB total JSON, including base64 images and history |
| Image formats | JPEG and PNG; base64 adds roughly one third to file size |
| Rate and concurrency | No fixed per-key allowance is published; contact support for capacity commitments |
| Serving deadlines | Up to 10 minutes for prefill; 15 minutes for a non-streaming response. Client/network timeouts may be shorter. |
Set a client timeout explicitly, such as timeout=600.0 on OpenAI(...),
and use streaming for long outputs. Incomplete streams cannot be resumed;
retries are new requests and may incur additional charges. Do not replay
completed tool actions.
Errors
| Status | Action |
|---|---|
400 / 422 | Fix the request or unsupported parameter. |
401 / 403 | Check the key and account/model access. |
402 | Check your Console balance. |
404 | Check the model ID and endpoint. |
413 | Reduce request size. |
429 / 503 | Reduce concurrency; back off with jitter and respect Retry-After. |
500 / 502 / 504 | Check service status; retry transient failures with a retry limit. |
Support
Service status · support@river.ai
Include the model, endpoint, request time/timezone, status code, and
x-request-id when available. Never include your API key.
Data handling
Default Commercial terms: no training on
your inputs or outputs; User Content retained up to 2 days after the session
ends, with legal, safety/security/compliance, and agreed exceptions.
De-identified aggregate statistics may be retained longer. Testing terms and signed addenda
may differ; store: false is not a zero-data-retention agreement.