Rate limits
Per-minute burst, concurrent stream cap, and per-key budget.
Cocobox applies three independent limits to every API key.
1. Burst (per minute)
50 requests per minute, per key.
The window is a 60-second sliding window. Once you cross 50 requests in any 60 seconds, the next requests return 429 rate_limited with a Retry-After header in seconds.
2. Concurrent streams
5 concurrent streaming requests, per key.
Streaming requests (/v1/chat/completions with stream:true, /v1/messages with stream:true, MCP responses) hold a long-lived connection. The 6th simultaneous stream is rejected with 429 rate_limited.
3. Per-key request budget
Optional, set when you create the key.
If you create a key with requestBudget: 10000, the 10,001st request returns 429 budget_exhausted. To raise or remove the cap, PATCH /api-keys/:id with the new value.
Headers
Every successful response includes:
X-RateLimit-Limit: 50
X-RateLimit-Remaining: 47
X-RateLimit-Reset: 1735689660
X-RateLimit-Reset is the Unix epoch (seconds) at which the burst window will fully refresh.
For per-key budget, when present:
X-Budget-Limit: 10000
X-Budget-Remaining: 9682
Best practices
- Back off exponentially with jitter. Doubling intervals starting at 1s, with ±20% jitter, works well.
- Honor
Retry-After. Never retry sooner. - Use a per-machine key. Three engineers sharing one key will share one bucket.
- Cache on your side. Identical prompts produce billable provider calls every time — the API does not cache responses for you.
Need higher limits?
Contact devs@cocobox.io with your workspace id and a description of the workload. We raise burst limits case-by-case for legitimate usage.