Rate limits and idempotency
The per-credential limit, and the header every write requires so a retry never creates a second row.
Rate limits
Limits are per credential, per rolling minute.
| Credential | Requests per minute |
|---|---|
| Test | 1200 |
| Live | 600 |
Exceeding one answers 429 rate_limited with Retry-After: 60. Wait the named number of seconds
and try again. Both SDKs honour the header for you.
Failed authentication is limited separately, 20 attempts per minute per source address, so guessing a key is throttled without the limit touching a working integration.
If your workload needs more, ask before you build around it. Batching a backfill behind a queue is
usually the answer, and starting_after paging is cheaper than parallel requests.
Idempotency
Every write requires an Idempotency-Key header. That is every POST, PATCH and DELETE
except POST /v1/realtime/tokens, which mints a short-lived secret rather than a resource and so
has nothing to replay. A write without one answers 400 invalid_request.
The key is yours to choose. It must be 16 to 128 printable ASCII characters and it must be unguessable enough not to collide with another attempt of your own.
curl https://api.cuvo.co/v1/patients \
-H "Authorization: Bearer cuvo_sk_test_…" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: patient-crm-88213-attempt-1" \
-d '{ … }'
What the key buys you
Retry a request that timed out, with the same key and the same body, and you get the first attempt's
answer back rather than a second patient, case or message. The replay carries
Idempotent-Replayed: true.
This is the only safe way to retry a write. A write with no key cannot be retried, because a timeout and a success are indistinguishable from your side.
The rules
- Same key, same request: replay. The stored status and body are returned verbatim.
- Same key, different request:
409 idempotency_conflict. The method and the path are part of the comparison, so one key reused against another route is caught too. - The key is claimed before validation. A retry of a request that failed validation replays that
same
400rather than being judged fresh each time. - A stored answer lives 24 hours. After that the key is free and the request is treated as new.
- 401, 403 and 429 are never stored. Those are decided before the handler runs, and replaying one would pin you to a credential problem you have since fixed.
- 5xx is never stored. A retry after a fault has to actually retry.
- A concurrent retry does not race. Whichever request records its answer first wins, and the other replays it.
Choosing keys
Derive the key from the thing you are creating, not from the moment you tried.
patient:<your external id>
case:<your order id>:v1
message:<thread id>:<sequence>
A key like uuid() generated inside the retry loop defeats the whole mechanism: every attempt is a
new key, so every attempt is a new row. Generate it once, alongside the job, and reuse it for every
attempt of that job.
Both SDKs mint a key per call and reuse it across their own retries. Pass your own when the retry comes from your queue rather than from theirs, which is the case that matters.
Body size
Request bodies are capped at 1 MB, answered with 413 payload_too_large. Files never go through a
request body; see Files.
Being a good client
- Page with
starting_afterrather than issuing parallel requests for a range. - Poll
GET /v1/events?since=…on a sensible interval instead of polling each case. - Prefer a webhook endpoint to polling at all. It costs you nothing against the limit.
- Back off on 429 and on 5xx, with jitter, so a fleet of your workers does not retry in lockstep.