Skip to content

Rate limits and idempotency

The per-credential limit, and the header every write requires so a retry never creates a second row.

Rate limits

Limits are per credential, per rolling minute.

CredentialRequests per minute
Test1200
Live600

Exceeding one answers 429 rate_limited with Retry-After: 60. Wait the named number of seconds and try again. Both SDKs honour the header for you.

Failed authentication is limited separately, 20 attempts per minute per source address, so guessing a key is throttled without the limit touching a working integration.

If your workload needs more, ask before you build around it. Batching a backfill behind a queue is usually the answer, and starting_after paging is cheaper than parallel requests.

Idempotency

Every write requires an Idempotency-Key header. That is every POST, PATCH and DELETE except POST /v1/realtime/tokens, which mints a short-lived secret rather than a resource and so has nothing to replay. A write without one answers 400 invalid_request.

The key is yours to choose. It must be 16 to 128 printable ASCII characters and it must be unguessable enough not to collide with another attempt of your own.

curl https://api.cuvo.co/v1/patients \
  -H "Authorization: Bearer cuvo_sk_test_…" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: patient-crm-88213-attempt-1" \
  -d '{ … }'

What the key buys you

Retry a request that timed out, with the same key and the same body, and you get the first attempt's answer back rather than a second patient, case or message. The replay carries Idempotent-Replayed: true.

This is the only safe way to retry a write. A write with no key cannot be retried, because a timeout and a success are indistinguishable from your side.

The rules

  • Same key, same request: replay. The stored status and body are returned verbatim.
  • Same key, different request: 409 idempotency_conflict. The method and the path are part of the comparison, so one key reused against another route is caught too.
  • The key is claimed before validation. A retry of a request that failed validation replays that same 400 rather than being judged fresh each time.
  • A stored answer lives 24 hours. After that the key is free and the request is treated as new.
  • 401, 403 and 429 are never stored. Those are decided before the handler runs, and replaying one would pin you to a credential problem you have since fixed.
  • 5xx is never stored. A retry after a fault has to actually retry.
  • A concurrent retry does not race. Whichever request records its answer first wins, and the other replays it.

Choosing keys

Derive the key from the thing you are creating, not from the moment you tried.

patient:<your external id>
case:<your order id>:v1
message:<thread id>:<sequence>

A key like uuid() generated inside the retry loop defeats the whole mechanism: every attempt is a new key, so every attempt is a new row. Generate it once, alongside the job, and reuse it for every attempt of that job.

Both SDKs mint a key per call and reuse it across their own retries. Pass your own when the retry comes from your queue rather than from theirs, which is the case that matters.

Body size

Request bodies are capped at 1 MB, answered with 413 payload_too_large. Files never go through a request body; see Files.

Being a good client

  • Page with starting_after rather than issuing parallel requests for a range.
  • Poll GET /v1/events?since=… on a sensible interval instead of polling each case.
  • Prefer a webhook endpoint to polling at all. It costs you nothing against the limit.
  • Back off on 429 and on 5xx, with jitter, so a fleet of your workers does not retry in lockstep.