Rate limits

Per-minute burst, concurrent stream cap, and per-key budget.

Last updated

Cocobox applies three independent limits to every API key.

1. Burst (per minute)

50 requests per minute, per key.

The window is a 60-second sliding window. Once you cross 50 requests in any 60 seconds, the next requests return 429 rate_limited with a Retry-After header in seconds.

2. Concurrent streams

5 concurrent streaming requests, per key.

Streaming requests (/v1/chat/completions with stream:true, /v1/messages with stream:true, MCP responses) hold a long-lived connection. The 6th simultaneous stream is rejected with 429 rate_limited.

3. Per-key request budget

Optional, set when you create the key.

If you create a key with requestBudget: 10000, the 10,001st request returns 429 budget_exhausted. To raise or remove the cap, PATCH /api-keys/:id with the new value.

Headers

Every successful response includes:

X-RateLimit-Limit: 50
X-RateLimit-Remaining: 47
X-RateLimit-Reset: 1735689660

X-RateLimit-Reset is the Unix epoch (seconds) at which the burst window will fully refresh.

For per-key budget, when present:

X-Budget-Limit: 10000
X-Budget-Remaining: 9682

Best practices

  • Back off exponentially with jitter. Doubling intervals starting at 1s, with ±20% jitter, works well.
  • Honor Retry-After. Never retry sooner.
  • Use a per-machine key. Three engineers sharing one key will share one bucket.
  • Cache on your side. Identical prompts produce billable provider calls every time — the API does not cache responses for you.

Need higher limits?

Contact devs@cocobox.io with your workspace id and a description of the workload. We raise burst limits case-by-case for legitimate usage.